The gist
In 76 years, AI has gone from Alan Turing's dream to an assistant that understands everyday language, writes working code and reasons through hard problems. It didn't happen "out of nowhere" in 2022 with ChatGPT. It's the result of 30+ key events, 5 ways of thinking about the problem, and 3 ice ages when the whole field froze.
In the History of AI lesson we skimmed the surface: the 5 biggest milestones. Here we go deeper. You'll learn why the era of agents started right now, in 2024-2026. Why Claude came from a separate company. Which people built this industry. And where we're headed next.
This isn't a history textbook. It's a map of the terrain. To understand what you're doing today and why, you need to know where it all came from.
The main rule of this lesson
No revolution in AI was a surprise. Every "suddenly" is 10-20 years of quiet work by a small group of scientists. ImageNet (2012) had been building since 2009. The Transformer (2017) grew out of the attention mechanisms of 2014-2015. ChatGPT (2022) came out of GPT-3 (2020), which came out of BERT (2018). When you see a fresh "breakthrough," look for the people who were working on it 5 years earlier.
The 5 eras of AI: a map of the terrain
Era 1: 1950-1956 — The dream of a thinking machine Era 2: 1956-1974 — Symbolic AI and early hope Era 3: 1974-2010 — Machine learning and two winters Era 4: 2012-2017 — The Deep Learning Revolution Era 5: 2018-2026 — LLMs and the era of agents
Next, each era in detail. With years, people and events. You don't need to memorize it. You need to understand the logic of how one led to the next.
Era 1: The dream of a thinking machine (1950-1956)
The big question of the era: can a machine think?
1950: Alan Turing publishes "Computing Machinery and Intelligence"
By this point Turing had already helped break the Enigma code and laid the theoretical foundation of computer science (the "Turing machine," 1936). In his paper he asks a simple question: instead of "can a machine think" (a philosophical question), ask "can a machine imitate a human so well that the other person can't tell the difference" (a question you can test).
This is the Imitation Game, later called the Turing Test. A machine passes if a human judge, after 5 minutes of written conversation, can't tell which of the two participants is the machine.
Turing predicted that by 2000 machines would pass the test "70% of the time." Reality: in a 2024 experiment at the University of California, San Diego, people mistook GPT-4 for a human in 54% of five-minute conversations, a borderline result.
1951: Marvin Minsky builds SNARC
Minsky, a 24-year-old Harvard student, builds SNARC (Stochastic Neural Analog Reinforcement Calculator), the first neural network in hardware. 40 neurons made of vacuum tubes, plus electric motors and telephone relays. The machine learned to find its way out of a maze through positive reinforcement.
It's the first physical version of the idea of "neurons in hardware," which would become the basis of all deep learning 60 years later.
1956: The Dartmouth Conference, the birth of AI
In June 1956, an 8-week summer workshop takes place at Dartmouth College (New Hampshire, USA). The organizers: John McCarthy, Marvin Minsky, Claude Shannon, Nathaniel Rochester. The conference budget: $7,500.
In the grant proposal, McCarthy uses the term "Artificial Intelligence" for the first time. Not "thinking machines," not "cybernetics," but AI. The term sticks.
The conference's forecast: "in 20 years machines will be able to do any work a human can do." Reality: it took 70+ years.
Era 2: Symbolic AI and early hope (1956-1974)
The big idea of the era: AI = logic + rules. Describe enough rules and the machine will think.
1957: Frank Rosenblatt invents the Perceptron
Rosenblatt, a psychologist at Cornell, builds the Mark I Perceptron, a machine that learns to recognize letters of the alphabet. It's a linear classifier: inputs → weights → a threshold decision.
In 1958 the New York Times reported that the US Navy had revealed an electronic machine it expected to walk, talk, see, write, reproduce itself and be conscious of its existence. The hype went through the roof.
This was the first AI hype cycle. Rosenblatt predicted machines would overtake humans within a few years.
1965: Joseph Weizenbaum creates ELIZA
At MIT, Weizenbaum writes ELIZA, the first chatbot. It imitates a therapist from the Rogerian school. A simple program: it spots keywords and turns the user's sentence back into a question.
User: I'm worried about my mother. ELIZA: Tell me about your mother. User: She always criticizes me. ELIZA: Who else always criticizes you?
The shocking discovery: people started talking to ELIZA about real problems. Weizenbaum's secretary asked him to leave the room so she could talk to ELIZA in private. This is the ELIZA effect: people credit a machine with understanding it doesn't have.
The lesson for 2026: people still credit LLMs with understanding they don't have. The ELIZA effect isn't a bug from 1965. It's a feature of human perception.
1969: Minsky and Papert publish "Perceptrons"
Minsky and Seymour Papert publish the book "Perceptrons," a formal mathematical analysis of Rosenblatt's perceptron. They prove that a single-layer perceptron can't learn the XOR function (exclusive OR).
A multi-layer perceptron can, but in 1969 there was no efficient way to train one. The book was taken as a death sentence for neural networks.
The consequences:
- Funding for neural networks drops sharply
- Rosenblatt dies in 1971 (a boating accident, at 43)
- The field of neural networks freezes for 17 years, until backpropagation (1986)
The 1970s: Expert systems
Alongside neural networks, symbolic AI develops: explicit rules + logic.
MYCIN (Stanford, 1972): diagnosing bacterial infections. 600 rules like "if symptom X and test result Y, then the probability of bacterium Z = 0.7." Its diagnostic accuracy was higher than that of the average junior doctor.
DENDRAL (Stanford): analyzing the mass spectra of chemical compounds.
Picture this: alchemy. Lots of work, lots of enthusiasm, limited results. The hypothesis "describe all the rules and the machine will be smarter than a human" didn't stand the test of time: there turned out to be too many rules, and they contradicted each other.
Era 3: Machine learning and two winters (1974-2010)
The big idea of the era: AI = find patterns in data instead of writing rules by hand.
1974-1980: The first AI winter
In the UK, the Lighthill Report (1973) comes out, a review of the state of AI commissioned by the British government. The verdict: inflated promises weren't kept, there's no progress, cut the funding.
The consequences were global:
- DARPA (US) sharply cuts AI funding
- AI labs close on both sides of the Atlantic
- The term "AI" becomes toxic: scientists rename their projects "knowledge-based systems" or "pattern recognition"
This is the first AI winter (1974-1980). It lasted 6 years.
1980-1987: A brief thaw
Expert systems have a commercial boom. LISP machines (specialized hardware for the LISP language) become an industry of their own. Companies: Symbolics, LMI, TI Explorer.
Japan's "Fifth Generation Computing" project (1982): $850 million over 10 years, aiming for parallel computers with a natural language interface.
1986: The backpropagation revival
Geoffrey Hinton, David Rumelhart and Ronald Williams publish the paper "Learning representations by back-propagating errors" in Nature. The algorithm itself was already known (Werbos 1974, Linnainmaa 1970), but they show how to train multi-layer neural networks efficiently.
Technically, this solves the XOR problem from 1969. But there are no GPUs yet, no data, and the whole world is caught up in expert systems.
1987-1993: The second AI winter
LISP machines collapse. Symbolics goes bankrupt. Japan's Fifth Generation project is declared a failure. Expert systems turn out to be brittle: they work on narrow tasks and don't scale.
The second AI winter was shorter than the first, but more painful for the commercial sector.
1997: Deep Blue beats Kasparov
On May 11, 1997, IBM's Deep Blue wins a match against world champion Garry Kasparov, 3.5 to 2.5. Chess, a game considered the peak of human intelligence, had fallen.
An important caveat: Deep Blue didn't "learn" in the modern sense. It was brute force search plus hand-written rules of thumb from grandmasters. 200 million positions per second. It was a win for compute, not for AI.
The lesson: at that point, "winning at chess" and "understanding" were different tasks. That would become clear 30 years later.
2006: Hinton revives deep learning
Geoffrey Hinton publishes "A fast learning algorithm for deep belief nets." He shows how to train networks with many layers (before that, it worked poorly because of vanishing gradients).
The term "deep learning" enters common use. Hinton, based in Toronto, along with Yoshua Bengio (Montreal) and Yann LeCun (NYU), forms what's later called the "Canadian Mafia" of the coming AI revolution.
2009: Fei-Fei Li launches ImageNet
Fei-Fei Li launches the ImageNet project at Stanford: 14 million images labeled by hand across 22,000 categories. The work was done through Amazon Mechanical Turk, millions of micro-tasks for human labelers.
In 2009 nobody cared. In 2010 the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) launches, an annual competition. In 2010 and 2011 classical methods win, with error rates of 26-28%.
In 2012, everything blows up.
Era 4: The Deep Learning Revolution (2012-2017)
The big idea of the era: AI = multi-layer neural networks + GPUs + lots of data.
September 2012: AlexNet wins ImageNet
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton (all from Hinton's lab at the University of Toronto) enter AlexNet in the ILSVRC competition: a convolutional neural network with 8 layers, trained on 2 NVIDIA GTX 580 cards (3 GB of memory each, about $500).
The result: an error rate of 15.3%, against 26.2% for the second-place entry. A gap of almost 11 percentage points, unheard of in a field where a 1% improvement was normal.
This is the moment AI became "deep learning." Every ImageNet winner after that used neural networks. Classical computer vision methods (SIFT, HOG, SVM) died off within a year.
Hinton, Sutskever and Krizhevsky founded DNNresearch, a startup sold to Google in 2013 for $44 million (Hinton went to Google Brain, and Sutskever later to OpenAI).
2014: GANs (Goodfellow)
Ian Goodfellow comes up with Generative Adversarial Networks during a conversation at a bar in Montreal. The idea: two networks play a game. The generator creates images, the discriminator tells real ones from generated ones. Both get better.
Five years later, GANs were creating deepfakes you can't tell from reality.
2015: AlphaGo (DeepMind)
DeepMind (London, founded in 2010 by Demis Hassabis, Shane Legg and Mustafa Suleyman, bought by Google in 2014 for $500M) releases AlphaGo.
Go was considered impossible for AI. A chess position has about 35 possible moves; Go has about 250. The search tree is exponentially bigger. Experts said "AI will beat Go in 30+ years."
AlphaGo uses deep learning + Monte Carlo Tree Search + reinforcement learning. It trained on 30 million positions from professional games.
In October 2015 it beats European champion Fan Hui 5-0. The first program to beat a professional Go player without a handicap.
March 2016: AlphaGo beats Lee Sedol
In Seoul, in a 5-game match, AlphaGo beats Lee Sedol (one of the strongest players in the world over the previous decade) 4-1. Every game was streamed live, with millions of viewers across Asia.
Move 37 in the second game was a move human professionals hadn't considered. Not an "algorithm mistake": it turned out to be a brilliant move. After the match, Lee Sedol said AlphaGo had shown him new ways to play Go.
This is the moment AI became something new in the public mind: not a "robot vacuum," but something that could surprise experts.
June 2017: "Attention Is All You Need"
Eight researchers from Google Brain and Google Research publish the paper "Attention Is All You Need." The authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, Illia Polosukhin.
They introduce the Transformer architecture, a neural network built on the self-attention mechanism. Before that, language processing used RNNs/LSTMs, sequential networks that process words one after another.
The Transformer processes the whole sequence in parallel. It trains 10-100 times faster on GPUs. And the quality is higher.
This paper has now been cited 100,000+ times. It's the foundation of every LLM: BERT, GPT, Claude, Gemini are all built on the Transformer.
A fun fact: most of the authors left Google within a few years, and many founded their own companies. It shows how talent migrated from Big Tech to AI startups.
Era 5: LLMs and the era of agents (2018-2026)
The big idea of the era: AI = one big model trained on the internet + tools + reasoning.
2018: BERT (Google)
Google releases BERT (Bidirectional Encoder Representations from Transformers), a Transformer trained to understand text in both directions (left to right and right to left). A revolution for Google Search: the system now understands queries better.
2018-2019: GPT-1, GPT-2 (OpenAI)
OpenAI (founded in 2015 by Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever and others; originally a nonprofit lab) releases GPT-1 (June 2018) and GPT-2 (February 2019).
GPT (Generative Pre-trained Transformer) is a Transformer trained to generate text. Pre-training on a huge body of internet text, then fine-tuning on specific tasks.
GPT-2 (1.5B parameters) generated text so well that OpenAI declared it "too dangerous for a full release." They released only a small version at first. Hype and criticism followed: some saw it as a marketing move.
May 2020: GPT-3
OpenAI releases GPT-3: 175 billion parameters. 100 times bigger than GPT-2. By estimates, training cost several million dollars in compute.
GPT-3's big discovery was in-context learning. Show the model 2-3 examples of a task in the prompt, and it learns to handle new cases. No fine-tuning, no retraining. A single prompt can teach it a new task.
This changed everything. Before GPT-3, every task needed its own model. After: one model + the right prompt = any task.
API access through the playground opened in June 2020. Developers around the world started building products on GPT-3. Prompt engineering appeared as a profession.
2021: Anthropic is founded
At the end of 2020, Dario Amodei (VP of Research at OpenAI) and his sister Daniela Amodei (VP of Operations) leave OpenAI. Several senior researchers leave with them: Tom Brown (lead author of the GPT-3 paper), Sam McCandlish, Jack Clark, Jared Kaplan and others.
The reason: disagreements with OpenAI about the priority of AI safety. They felt commercialization was outpacing safety.
In 2021 they found Anthropic in San Francisco. A Public Benefit Corporation. The original focus: research on AI safety. Starting round of $124M (then $580M in 2022).
2022: Constitutional AI
Anthropic publishes the paper "Constitutional AI: Harmlessness from AI Feedback." The idea: instead of RLHF (Reinforcement Learning from Human Feedback), where people rate every answer the model gives, use a set of principles (a constitution) and have the AI itself evaluate its answers.
This makes safety training scalable. And it becomes Claude's approach.
November 2022: ChatGPT launches
On November 30, 2022, OpenAI releases ChatGPT, a chat interface for GPT-3.5. Free, no waitlist, available to everyone.
Within 5 days: 1 million users. Within 2 months: 100 million. The fastest-growing product in history at the time. For comparison: TikTok reached 100M in 9 months, Instagram in 2.5 years.
ChatGPT was the moment "AI" stopped being a topic for the technical community. It became a topic for CNN, the BBC, and parents calling their kids to ask "what is this thing."
March 2023: GPT-4 + Claude goes public
OpenAI releases GPT-4. Anthropic comes out of research-only mode and publicly releases Claude (first through a partnership with Slack, then through claude.ai).
Claude 1 was a competitor to GPT-3.5. Not better, but safer (Constitutional AI was working).
July 2023: Claude 2 + 100K context
Anthropic releases Claude 2, better than Claude 1 and a competitor to GPT-4 on some tasks.
The main thing: a context window of 100,000 tokens. For the first time, an AI could hold an entire book (300 pages) in memory in a single request. Before that, models had 2K-8K of context.
This changed how people work with documents. You could load a whole codebase, a whole book, a whole stack of legal documents, and the model sees all of it.
March 2024: The Claude 3 family
Anthropic releases Claude 3 in three sizes:
- Haiku: fastest, cheapest
- Sonnet: balanced
- Opus: most capable
Claude 3 Opus outperformed GPT-4 on a number of benchmarks, the first time a non-OpenAI model took the lead. A real alternative appeared on the market.
June 2024: Claude 3.5 Sonnet
Three months after Claude 3 came Claude 3.5 Sonnet. Better than Claude 3 Opus, yet cheaper and faster. Especially strong at programming: 92% on the HumanEval benchmark.
This was the smartphone moment for programming: when 3.5 Sonnet came out, many developers switched to Claude for code.
October 2024: Computer Use
Anthropic releases Computer Use: Claude can now operate a computer. It sees the screen, moves the mouse, types text. Not perfectly, but it works.
This is the beginning of the era of agents. AI stops being a chat interface and becomes "the one who does the work."
November 2024: MCP (Model Context Protocol)
Anthropic releases the Model Context Protocol, an open standard for connecting tools to LLMs. Not a proprietary API, not a closed plugin format, but an open specification anyone can implement.
MCP gives LLMs universal outlets for plugging in outside systems: the file system, GitHub, databases, APIs. One protocol, any tool.
It becomes an industry standard. Within 12 months: hundreds of MCP servers for everything.
2025-2026: A cascade of Claude releases
In 2025-2026, new versions of Claude come out every one to three months. Key milestones:
- Claude 3.7 Sonnet (February 2025): the first hybrid model with a reasoning mode
- Claude 4 (May 2025) and Claude 4.5 (fall 2025)
- Claude Opus 4.7 (April 2026)
- Claude Fable 5 (June 2026), Sonnet 5 (June 2026), Opus 5 (July 2026)
- Fable 5.1 (September 1, 2026), Opus 5.5 (September 22, 2026), Sonnet 5.5 (September 28, 2026)
Each one is a noticeable step forward. A reasoning mode (the model "thinks" before answering, similar to OpenAI's o-series models), context of up to 1M tokens, better work with tools. For the current list of models, see the What's current page.
2025: Claude Code
Alongside the models, Anthropic releases Claude Code (first version in February 2025), a terminal-native AI assistant. Not a chat window, but a command-line tool that lives in your terminal and works with your codebase.
It changes how developers work. You sit in the terminal, type claude, describe the task, and the agent writes code into files, runs tests and makes commits. As of October 2026, Claude Code also works in VS Code, JetBrains, the Claude desktop app, the browser and on your phone.
2026 (now): The agent ecosystem
October 2026: we're at the moment when the agent ecosystem is taking shape:
- Frameworks: LangGraph, CrewAI, Mastra, AutoGen and others
- Stack: Skills + Hooks + Subagents + Plugins for Claude Code
- Local AI: nanoClaude, OpenClaw, Hermes, small models that run on a laptop
Production-grade agents are already at work: automating customer support, content marketing, sales outreach, code review. But it's early morning. Most of the work is ahead.
5 AI paradigms: how the approach changed
| Paradigm | Era | Idea | Example | What killed it |
|---|---|---|---|---|
| Symbolic AI | 1956-1980 | AI = logic + rules | MYCIN, DENDRAL | Brittleness, combinatorial explosion |
| Statistical ML | 1990-2010 | AI = find patterns in data | SVM, Random Forests, Naive Bayes | Weak on complex tasks |
| Deep Learning | 2012-2018 | Multi-layer neural networks + GPUs | AlexNet, AlphaGo, ResNet | Not killed; grew into LLMs |
| Pretrained LLMs | 2018-2023 | One big model trained on the internet | GPT-3, Claude 2, BERT | Not killed; enriched with agency |
| Agentic AI | 2024-now | Model + tools + memory + reasoning | Claude Code, Computer Use, agents | Active era, outcome unknown |
3 AI winters: what went wrong
| Winter | Years | Cause | Result |
|---|---|---|---|
| First | 1974-1980 | Lighthill Report, early promises not kept | Funding drops sharply, the field freezes |
| Second | 1987-1993 | Collapse of LISP machines, brittle expert systems | The commercial sector leaves for 10 years |
| Third? | Not yet | Possible risks: capability plateau, regulation, running out of training data | Watch closely |
Could there be a third winter?
Arguments for "yes":
- Capability plateau: some experts think improvements are slowing down; the gap between neighboring model generations is smaller than it was between GPT-3 and GPT-4
- Training data exhaustion: we've used up the "good" internet for training
- Regulation: the EU AI Act and various national regulations could slow deployment
- Economic reality: many startups have weak unit economics. The bubble could burst
Arguments for "no":
- Real productivity: AI already delivers measurable returns in coding, support and content
- Hardware progress: NVIDIA keeps releasing better and better GPUs
- The agentic frontier: the shift to agents has only just begun, and the potential is huge
- Capital: huge sums have been invested in the industry, and the ecosystem won't let it die quickly
The real risk in the next few years isn't a "winter for all of AI" but a correction: some AI startups won't survive, and the ones with a real product will remain. That's normal and healthy.
Why Claude (and not another model)
In this course, the main track runs through Claude. Not because Claude is "the best at everything," but because:
Technical reasons
- Constitutional AI: training through safety principles. The goal is more predictable behavior (but attacks on models still happen; see the Prompt Injection Defense lesson).
- Long context first: Claude 2 supported 100K in 2023, half a year ahead of OpenAI. Now it's 1M tokens.
- Code-native focus: Claude models do well at programming. Check current independent coding benchmarks: the rankings change with every release.
- Computer Use first: Anthropic released Computer Use in October 2024, ahead of everyone else.
- The MCP standard: an open protocol, not a proprietary one. Anthropic isn't trying to lock in the ecosystem.
Business reasons
- Ex-OpenAI founders: Dario Amodei (former VP of Research at OpenAI) and his team are the same people who built GPT-2/GPT-3. They know what they're doing.
- A long-term focus on safety: Anthropic builds its brand around the safety and predictability of its models.
- Documentation: Anthropic publishes detailed migration guides between model generations. That said, a generation change sometimes breaks old settings: for example, the newest models removed manual sampling parameters (temperature and the like).
Where Claude isn't the leader
- Multimodal video: Gemini accepts video as input; Claude takes text and images
- Real-time speech: ChatGPT's voice modes and OpenAI's Realtime API
- Image generation: Midjourney, FLUX, and the built-in generation in ChatGPT and Gemini are better for pictures
- Open source: there are no open weights for Claude (open weights exist for DeepSeek, Qwen, some Mistral models and earlier Llama releases)
Not a silver bullet. In this course Claude is the foundation, but we'll point out when another model is a better fit.
What happened between 2023 and 2026: the main shifts
| Parameter | 2023 | 2026 | How much it changed |
|---|---|---|---|
| Context window (how much text the model keeps in its head) | 4-8 thousand tokens | 1 million tokens and up | 125-250 times more |
| What it understands | Text only | Text, images, audio and video | 4 times as many kinds of data |
| Tools | None | Function calling → the shared MCP standard | didn't exist before |
| Reasoning | Only through cleverly written prompts | Step-by-step reasoning → reasoning mode (adaptive thinking) | a qualitative leap |
| Price of input text | $30 per 1M tokens (GPT-4, 2023) | $1 per 1M tokens (Claude Haiku 4.5, as of October 2026) | 30 times cheaper (models of different classes) |
| Speed of a simple answer | 30 seconds | 1 second | 30 times faster |
| Independence | Answers in a chat | Multi-step plans the model carries out on its own | a qualitative leap |
| Running on your own computer | Cloud only | Open models on a laptop | now possible |
The most important shift is agency. Not "AI answers," but "AI does." This changes the human's role: from "tool operator" to "manager of a team of agents."
Key figures: short profiles
Alan Turing (1912-1954)
British mathematician. The father of computer science thinking. Helped break the Enigma code in World War II. Gay at a time when that was a crime in the UK. In 1952 he was arrested, convicted and subjected to chemical castration. In 1954 he died of cyanide poisoning (officially a suicide). He was pardoned only in 2013, posthumously. Without Turing, computer science as a field would not exist.
John McCarthy (1927-2011)
American computer scientist. Named the field "Artificial Intelligence" in 1956. Created the LISP programming language (1958), which all early AI systems were written in. Founded the Stanford AI Lab (1963). 1971: Turing Award.
Marvin Minsky (1927-2016)
MIT. Co-founder of the MIT AI Lab (1959). Built SNARC (1951). Co-author of "Perceptrons" (1969). Influential, but at times controversial: his critique of neural networks held back their development for 17 years. 1969: Turing Award.
Geoffrey Hinton (born 1947)
"The Godfather of deep learning." University of Toronto. The backpropagation revival (1986), deep belief networks (2006), AlexNet (2012, through his student Krizhevsky). Google Brain 2013-2023. In 2023 he left Google so he could speak freely about the risks of AI. 2018 Turing Award (with LeCun and Bengio) and the 2024 Nobel Prize in Physics (for his work on neural networks). He now warns about the risks of AGI.
Yann LeCun (born 1960)
French computer scientist. NYU + Meta (Chief AI Scientist from 2013 until the end of 2025; then founded the startup AMI Labs). Pioneered convolutional neural networks (CNNs) in 1989 for reading handwritten digits, such as ZIP codes for the US Postal Service. 2018 Turing Award (with Hinton and Bengio). Skeptical of LLMs as a path to AGI; he's betting on "world models."
Yoshua Bengio (born 1964)
Université de Montréal. The third member of the "Canadian Mafia." 2018 Turing Award. Since 2023 he has focused on AI safety.
Ilya Sutskever (born 1986)
Russian-born Israeli-Canadian researcher. Hinton's student. Co-author of AlexNet (2012). Co-founder and Chief Scientist of OpenAI (2015-2024). Led the training of GPT-3 and GPT-4. In May 2024 he left OpenAI after the conflict around Sam Altman. Founded Safe Superintelligence Inc., whose only goal is to build safe AGI.
Dario Amodei (born around 1983)
American physicist. PhD from Princeton. VP of Research at OpenAI 2018-2020. Left OpenAI with his sister Daniela and in 2021 they founded Anthropic. CEO. Optimistic about AI progress, while consistently advocating for safety.
Sam Altman (born 1985)
American entrepreneur. President of Y Combinator 2014-2019. Co-founder of OpenAI in 2015. CEO of OpenAI since 2019. In November 2023 the board fired him, and within 4 days he was back (it became known as the "OpenAI drama"): the vast majority of employees threatened to quit, and the board reinstated him. One of the most influential figures in the AI industry.
Demis Hassabis (born 1976)
British neuroscientist and former game designer (Theme Park, 1994). Founder of DeepMind (2010, sold to Google in 2014). AlphaGo, AlphaFold (predicting protein structures). 2024 Nobel Prize in Chemistry for AlphaFold (with John Jumper).
Andrej Karpathy (born around 1987)
Slovak-born researcher. PhD from Stanford under Fei-Fei Li. Founding member of OpenAI (2015-2017). Director of AI at Tesla 2017-2022. Returned to OpenAI 2023-2024, then left. Now focused on education: open courses, YouTube tutorials. Huge influence on the community.
Fei-Fei Li (born 1976)
Chinese-American researcher. Stanford. Creator of ImageNet (2009). Co-Director of Stanford HAI (the Human-Centered AI Institute). AI4ALL, a nonprofit for diversity in AI. Author of the book "The Worlds I See" (2023).
Ian Goodfellow (born 1985)
Inventor of GANs (2014). Google → Apple → DeepMind. An influential figure, but he avoids the public spotlight.
2026 outlook: where we are now
Capability
- Frontier models: several companies release top-tier models (Claude Fable and Opus, GPT-6, Gemini 3.x, DeepSeek V4 and others). The rankings change with every release; for the current list, see the What's current page
- Reasoning models: a reasoning mode is built into the models of every major company and gives a steady improvement on hard tasks (math, code, science)
- Multimodal: all frontier models understand text + images, and many also audio + video
- Long context: 1M tokens on a number of models; exact values are in each provider's documentation
Adoption
- Knowledge workers and programmers: the share of AI users is growing fast. For fresh numbers, check the AI Index Report (Stanford) and survey reports (for example, the Stack Overflow Developer Survey) rather than this lesson: the figures go out of date within months
- Business: many companies have an AI strategy; fewer actually use AI
Agents
- Early morning. Production-grade agents work in a few areas: customer support, content marketing, sales outreach, code review.
- Limits: decisions in medicine and law stay with specialists, and long chains of actions need human oversight.
AGI timeline (estimates)
Estimates vary enormously: heads of AI companies talk about a few years or the next decade, some well-known researchers believe LLMs won't lead to AGI and new architectures are needed, and skeptics allow that AGI may never appear at all. We don't give specific years here: they change with every interview.
The truth: nobody knows for sure. Five-year forecasts in AI have historically been wrong in both directions.
Risks
- Misuse: deepfakes, generated disinformation, autonomous weapons
- Job displacement: the tasks of translators, copywriters and junior developers are changing fastest (more in the AI without fear lesson)
- Alignment: whether models do what we actually want, not just what we literally ask
- Concentration of power: a small number of companies control the frontier (Anthropic, OpenAI, Google and others, plus Chinese labs like DeepSeek and Alibaba)
This isn't the Terminator. These are real concerns that need real work.
What "the era of agents" means in 2026
2018-2022: The era of pre-trained LLMs - AI answers questions - One prompt = one answer - No memory, no tools, no actions 2022-2024: The era of ChatGPT / conversational AI - Multi-turn conversations - Simple tools (web search, code execution) - Everything in a chat interface 2024-now: The era of agents - AI plans and carries out multi-step tasks - Tools + memory + reasoning - Works in your environment (terminal, computer, browser) - Can carry out tasks while you're away
What this means for you
You're learning at a moment when the rules are being written right now. Best practices for prompt engineering settled in 2023-2024. Best practices for agent engineering are being written in 2026. In 2 years they'll be learned, written down and taught in courses. Right now, it's the frontier.
That means:
- Fewer ready-made solutions ("google how to do X" won't work)
- More experimenting and learning from your own experience
- More chances to understand things more deeply than others, but no guarantee that it will bring in income
Checklist: what you understood from this lesson
What's next
If you're just starting to learn about AI: → How an LLM works inside: technical depth without formulas
If you want hands-on practice: → Installation and setup: your first command in the terminal
If you're interested in local AI and privacy: → Local AI Agents: small models on your laptop
If you want a deep dive into model architecture: → How an LLM works inside: the transformer and the attention mechanism without the math
If you're interested in the AGI timeline and risks: → AI Ethics & Safety: AI safety and how not to trust AI blindly
Sources
Foundational papers (read the originals if you have time)
- Turing 1950 — "Computing Machinery and Intelligence" — https://academic.oup.com/mind/article/LIX/236/433/986238
- Rosenblatt 1958 — "The Perceptron" — Psychological Review
- Rumelhart, Hinton, Williams 1986 — Backpropagation — Nature
- LeCun 1989 — CNNs for digit recognition
- Vaswani et al. 2017 — "Attention Is All You Need" — https://arxiv.org/abs/1706.03762
- Brown et al. 2020 — GPT-3 paper "Language Models are Few-Shot Learners" — https://arxiv.org/abs/2005.14165
- Bai et al. 2022 (Anthropic) — Constitutional AI — https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback
Research labs
- Anthropic Research — https://www.anthropic.com/research
- OpenAI Research — https://openai.com/research
- DeepMind Publications — https://deepmind.google/research/publications/
- Google AI — https://ai.google/
Games and benchmarks
- AlphaGo paper — Nature 2016 — https://www.nature.com/articles/nature16961
- ImageNet — https://www.image-net.org/
- MMLU benchmark — https://github.com/hendrycks/test
- HumanEval (code) — https://github.com/openai/human-eval
Industry reports
- AI Index Report (Stanford, latest edition) — https://aiindex.stanford.edu/
- State of AI Report (Nathan Benaich) — https://www.stateof.ai/
Books (if you want a deep dive)
- "The Coming Wave" — Mustafa Suleyman (2023)
- "Life 3.0" — Max Tegmark (2017)
- "Superintelligence" — Nick Bostrom (2014)
- "Human Compatible" — Stuart Russell (2019)
- "The Worlds I See" — Fei-Fei Li (2023)
- "Genius Makers" — Cade Metz (2021) — the best history of AI up to 2020
YouTube (learning visually)
- 3Blue1Brown — the Neural Networks series (visualizing the math)
- Andrej Karpathy — "Neural Networks: Zero to Hero" — building GPT from scratch
- Two Minute Papers — weekly reviews of research papers
Key takeaways
No revolution in AI was a surprise. ChatGPT in 2022 came out of GPT-3 in 2020. GPT-3 came out of the Transformer in 2017. The Transformer came out of the deep learning revival of 2006. Every "suddenly" is 10-20 years of quiet work by small teams.
Technology moves in waves of paradigms. Symbolic AI (1956-1980) → Statistical ML (1990-2010) → Deep Learning (2012-2018) → LLMs (2018-2023) → Agents (2024+). Old paradigms don't die. They become infrastructure for the new ones.
The era of agents has only just begun. Production-grade agents work in a few areas, but it's early morning. The rules and best practices are being written right now, in 2026. People learning today are closer to the frontier.
Module 0 navigation: ← History of AI | How an LLM works →
The mark stays in this browser only and is never sent anywhere. My progress