When Did AI First Appear? Not 1956, but 13 Years Earlier
Published September 30, 2026
The name was coined at Dartmouth in 1956, but the idea was written down in 1943. AI's history shows a pattern that keeps repeating.

The term artificial intelligence was coined in the summer of 1956 at Dartmouth College, by ten people who spent eight weeks arguing about one sentence. But that was the name, not the beginning of the idea. The idea had already been written down thirteen years earlier, in 1943, by two scientists working on biological neural networks. The equations they derived from the human brain are still the foundation of nearly every language model in use as of 2026.
The short answer: AI was born as a field name in 1956, as an idea in 1943, and as a testable question in 1949. Each date points at something different, and conflating them is the source of most public confusion about the field's history.
Thirteen years before the name existed
In 1943, neurophysiologist Warren McCulloch and mathematician Walter Pitts published a paper titled A Logical Calculus of the Ideas Immanent in Nervous Activity in the journal Bulletin of Mathematical Biophysics. It was not a computer program. It was an argument that neurons could be modelled as logic switches.
The McCulloch-Pitts neuron model is very simple: every input is 0 or 1 and carries a weight, the weighted inputs are summed and compared against a threshold. If it clears the threshold, the neuron fires. There is no abstraction inside it, no symbolic representation, no language.
What matters is not the neuron but the claim attached to it. If every activity in the brain can be written as binary logic operations, then anything that can be computed can be executed by a machine. Calculating machines already existed by the 1940s. What did not yet exist was a bridge from the work of nerves to the work of silicon, and that paper stated it formally.
Six years later, Alan Turing turned the question "can machines think?" into a testable one. His paper Computing Machinery and Intelligence appeared in Mind in 1950, opening with one of the discipline's most famous sentences: I propose to consider the question "Can machines think?" He did not answer it. He replaced it with the imitation game, now the Turing test: one human interrogator, one human foil, and a machine. If the interrogator cannot reliably tell them apart, the machine counts as intelligent.
That definition is brilliant because it removes the metaphor. "Thinking" has never been defined to anyone's satisfaction, so Turing did not try. He moved the problem somewhere it could be answered: observable behaviour.
The summer of 1956: eight weeks and half the grant money requested
The proposal is dated 31 August 1955. John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon made an explicit funding request to the Rockefeller Foundation: We propose that a 2 month, 10 man study of artificial intelligence be carried out during the summer of 1956 at Dartmouth College in Hanover, New Hampshire. The typescript ran to seventeen pages, and copies are still archived at Dartmouth and Stanford.
They got half: the proposal asked for $14,000 and the foundation awarded $7,000. That sounds like very little, but the context makes it legible. In 1956 the world had a few hundred computers, each the size of a room, with less computing power than one processor in a modern smartphone.
Eight weeks. Ten people. The participant list contains most of the names that still matter in the history of AI: McCarthy (Dartmouth), Minsky (Harvard), Shannon (Bell Labs), Rochester (IBM), Newell (RAND), Simon (Carnegie), Samuel (IBM), Solomonoff (MIT), Selfridge (MIT), and More (Princeton).
What was decided there was not technology but the name. The four authors had already used the phrase artificial intelligence in their 1955 document, and nobody had ever written those two words together as one term. That is why 1956 counts as the birth date: not because something was discovered, but because something that already existed finally got a name people could argue about.
One detail from the Dartmouth archives is often missed. While the proposal was being written, McCarthy recorded a hope that by summer 1956 he would have a machine model fairly close to the stage of ordinary computer programming. Sixty years later the language models still have the same shape: a program that writes programs.
Overpromising, and the first winter
After 1956, optimism ran very fast, and often too far.
In 1957, psychologist Frank Rosenblatt built the Perceptron, a single-layer neural network whose weights could be learned: artificial neurons that adjust themselves every time the answer comes out wrong. The press did not take long. Rosenblatt predicted it would recognise people and address them by name, and translate speech directly from one language into another.
The verdict came from inside the field itself. In 1969, Marvin Minsky and Seymour Papert published Perceptrons, showing that single-layer networks are mathematically incapable of solving problems like XOR, which require computation that cannot be linearly separated. The critique was correct, and Minsky himself acknowledged in the 1988 edition that the limitation had been visible since the mid-1960s and was only waiting to be written down.
What took its place was expert systems: software with hundreds of human-written rules for narrow problems. MYCIN diagnosed bacterial infections with rules that outperformed resident doctors, and could not reason outside them.
The second winter, when computing became a commodity
By 1985 the world had spent more than a billion US dollars on AI, almost all of it on expert systems and the companies built around them. The most cited example is XCON from Carnegie Mellon, which translated DEC customers' VAX requirements into production configurations. Its first release condensed expert knowledge into 480 configuration rules, and was credited with cutting delivery time from months to weeks.
That example feels quaint now, and that is the point. Expert systems worked because human knowledge had been gathered into an executable form. When it changed, rules had to be rewritten by hand. When input fell outside the anticipated range, the system failed, with no way to learn from it.
The next disaster came from price, not technology. The LISP machine market collapsed in 1987. Companies such as Symbolics had built expensive workstations optimised for one language, while ordinary personal computers had become fast enough to run CLIPS or a simple Lisp interpreter at usable speed. A half-billion-dollar industry disappeared in a single year. Symbolics went bankrupt in 1991.
What actually brought the field back
The answer is not a new algorithm. It is datasets and hardware.
ImageNet was assembled in 2009 at Stanford: more than 15 million labelled images across roughly 22,000 categories. What mattered was not the image count but the deliberate labelling noise, averaging about 1.2 labels per image. That turned image classification from an exhausting guessing exercise into a problem with a number attached that participants could be compared on.
In 2012, Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton won with AlexNet, a convolutional network with five convolutional layers, three fully connected layers, roughly 60 million parameters and 650,000 units. Their result on ILSVRC-2012:
| Model | Top-5 error (test) |
|---|---|
| AlexNet (SuperVision, Toronto) | 15.3% |
| Runner-up (SIFT + Fisher Vectors) | 26.2% |
| Gap | 10.9 points |
That number is what turned the field. Eleven percentage points on the same task, same dataset, within one year. AlexNet introduced no genuinely new concept: convolutional networks had existed since LeNet in 1998, and dropout for years. What was new was the combination, plus GPUs trainable in reasonable time, plus a dataset large enough to force generalisation. Three things that were not independent.
The jump was not about AI becoming "intelligent". AlexNet solved classification over 1,000 categories: a tightly defined task with answers checkable one at a time. What made it matter afterwards was generalising to tasks never seen before, which remains the same problem today under a different name.
Seventy years later, measured in months
On 30 November 2022, OpenAI opened ChatGPT to the public. The growth numbers that followed became the most widely cited adoption record in technology:
Those figures come from different bodies and were never standardised. Some count registered users, some monthly actives, some visits. But the direction is not ambiguous: roughly two months, while comparably aged applications took years. Analysts at UBS called it an unprecedented growth rate for a consumer application.
What is interesting about that table is not the speed but what stayed the same across seventy years. ChatGPT introduced nothing new in principle: the Transformer dates to 2017, RLHF was common, the scale of data and compute was clear. One thing differed. The product shipped to the public without an intermediary, so raising money that normally takes a decade happened in a quarter.
Analysis and what it means: two AI winters and their mechanism
The pattern is not a coincidence. Read the history as one sequence and the two funding collapses have almost the same shape.
The first winter runs 1974 to 1980, the second 1987 to 1993, and the second lasts longer than most accounts suggest. The first was triggered by the Lighthill report of 1973, written at the request of the UK's Science Research Council. It assessed AI research across Britain and concluded nothing had produced the promised impact. The argument was simple, and just as easy to reverse.
What this story usually skips is what happened in the laboratories rather than the headlines. When the Lighthill report appeared in 1973, the fast-growing SIGART group had 1,241 members. By 1978, in what is usually treated as the darkest stretch of the first winter, membership had reached roughly 3,500. The community was not shrinking. It grew nearly threefold, faster in proportional terms than ACM over the same period.
Those two facts are not in tension. What collapsed was not the number of people but the credibility of grant applicants. The distinction matters because the implications differ. If research interest had collapsed, the next revival would need a generation. If the label in a funding proposal had collapsed, the revival could happen within a few years, once computing got cheaper.
The same shift happened twice:
| Winter | Trigger | Fact that inverts the story |
|---|---|---|
| First, 1974-1980 | Lighthill report 1973 in the UK; DARPA grant cuts in the US | SIGART membership rose from 1,241 to roughly 3,500 by 1978 |
| Second, 1987-1993 | LISP machine market collapsed in 1987; rule-based industry lost its buyers | Commodity computing with a C implementation replaced dedicated machines; inference cost fell sharply |
What can be read off that table is one rule that held both times: a winter happens when the falling cost and the fixed cost meet in the wrong place. In the 1970s computing had not become cheap enough to reach the problem space being claimed. In the 1980s computing had become cheap, but the cost of knowledge never fell, because every rule had to be written by a human. The 2012 system moved two costs at once: computing got cheaper through GPUs, knowledge got cheaper because labels were produced at scale.
The most useful implication for a reader in 2026 is right here. A third winter, if it comes, will almost certainly not come from technical failure, because the techniques in use work. More likely is a split in the cost structure: the inference cost paid per token by users is separate from the knowledge cost paid by companies preparing data, labels, and evaluations. Winters happen not because AI fails, but because the cost that determines the value moves.
The whole chronology on one page
- 1943 — McCulloch and Pitts model the neuron as binary logic.
- 1950 — Turing replaces "can machines think?" with the imitation game.
- 1955 — Dartmouth proposal, 31 August, requests $14,000.
- 1956 — Dartmouth Summer Research Project, eight weeks. The term artificial intelligence is used for the first time as a field name.
- 1957-1958 — Rosenblatt's Perceptron, with press predictions that overshoot badly.
- 1966 — ELIZA, the first conversational program to make users feel they were talking to a person.
- 1969 — Perceptrons by Minsky and Papert; neural network research slows for nearly two decades.
- 1973 — Lighthill report; Britain withdraws support from almost every AI centre.
- 1974-1980 — the first AI winter.
- 1986 — backpropagation is popularised, neural networks return.
- 1987 — the LISP machine market collapses within a year.
- 1987-1993 — the second AI winter.
- 2012 — AlexNet: 15.3% top-5 error against 26.2% on ILSVRC.
- 2017 — the Transformer, the architecture behind every large language model in use as of 2026.
- 2022 — ChatGPT ships to the public, 100 million users in about two months.
Related tools
Free browser tools that apply to this topic.
Share this article
Share to
Related articles

September 30, 2026
20 Countries Cap Social Media for Kids, Indonesia Moves First
Australia from December 2025, Indonesia from March 2026. At least 20 countries have similar rules, but only some are actually in force.

September 30, 2026
AI Bots Are Now 57% of Web Traffic, and Nobody Reports It
Cloudflare counts 57.5 percent of web traffic as bots. The number is not as simple as it looks, and here is how to read it.

September 29, 2026
G30S From a Historian's View: The Numbers Nobody Counted
Historian Anhar Gonggong reads G30S as the end of a 45-year ideological argument. What remains is not a verdict but a number nobody ever counted.



