Questions

Questions people ask about AI

Short, plain answers to common questions about AI chatbots and the models behind them. Each answer lists the sources it rests on, and every number from them is labeled, like everything else on this site.

Reference Sources checked to

How it works

What’s the difference between AI, machine learning, and a large language model?

In short:They nest inside each other. AI is the broad field, machine learning is one way to build it, and a large language model is one kind of machine-learning system: the kind behind chatbots.

Artificial intelligence is the broad field of making computers do tasks that usually take human intelligence, such as recognizing speech or translating text.

Machine learning is the part of AI where a computer learns patterns from examples instead of following rules people wrote by hand. Deep learning does this with neural networks: many layers of simple calculations, whose are adjusted during training.

A is a trained on huge amounts of text to predict the next , a word or a piece of one. Think of the word suggestions on a phone keyboard, scaled up enormously. A chatbot is an app built around one: it sends your conversation to the model and shows the reply. This site is one, and shows each step.

See a reply made, step by step

Where does a chatbot’s knowledge come from? Does it look things up?

In short:From patterns in the text it was trained on, which stops at a cutoff date. It looks things up only if the app gives it a search tool, and this site doesn't.

Training adjusts a model's learned numbers until it predicts its training text well. What it learned is stored in those numbers, not kept as documents it can open and check.

That training text ends at a , like a snapshot taken on a certain day. For gpt-6-luna, the model this site uses, OpenAI lists a knowledge cutoff of May 18, 2026Reference: OpenAI, GPT-6 Luna model page, retrieved 2026-09-30. Later events are unknown to the model unless the app adds them to the text it sends.

Some chat apps can search the web as a : the model asks for a search, and the results are added to the text it works from. This site offers no tools or search, so every reply here comes from what the model learned in training, plus the text this app sends.

Why does it sometimes state false things so confidently?

In short:Because it scores which wording is likely to come next, not whether it's true. A fluent sentence can be likely and still false. These errors are often called “hallucinations”.

Each token is picked from options scored by how likely they are to come next, given the text so far. It's like a phone's word suggestions: they offer what usually comes next, not what's true. Nothing in that step checks facts, so a wrong answer can come out as smoothly as a right one, in the same confident tone.

OpenAI's researchers argue that training and testing make this worse. Most tests count only right answers, so a model that guesses scores better than one that admits it isn't sure. Facts that appear rarely in training text, like a particular person's birthday, are especially hard to get right.

Admitting uncertainty helps. In one of OpenAI's tests, a model that declined to answer 52%Reference: OpenAI, “Why language models hallucinate”, retrieved 2026-09-30 of the questions gave wrong answers 26%Reference: OpenAI, “Why language models hallucinate”, retrieved 2026-09-30 of the time. A model that almost never declined got slightly more right (24%Reference: OpenAI, “Why language models hallucinate”, retrieved 2026-09-30 against 22%Reference: OpenAI, “Why language models hallucinate”, retrieved 2026-09-30), but gave wrong answers 75%Reference: OpenAI, “Why language models hallucinate”, retrieved 2026-09-30 of the time.

See a recorded wrong answer in the sample

Reference Sources

Why do I get different answers to the same question?

In short:Each token is a weighted random pick among the options the model scores, so the same question can go down a different path each time.

At every step, the model scores the options for the next token, and one is picked at random, weighted by those scores: a . It's like a raffle where likelier options hold more tickets. The likeliest option usually wins, but not always, and one different pick early on changes everything after it.

Apps can make the pick more or less random with a setting called . Lower values make the likeliest options likelier; higher values spread the chances out.

Even with the same settings, OpenAI doesn't guarantee that identical requests give identical replies.

See a weighted random pick

Does it learn from my conversations?

In short:Not while you chat: its learned numbers stay fixed. But some companies use saved conversations to train future models, depending on your settings.

Within a conversation, a chatbot's “memory” is the app sending the again with each new message, like handing over the whole transcript each time. The model itself doesn't change, the way a calculator doesn't change from being used.

Memory features in some chat apps save details from your chats and add them to the text sent with later chats. That's still context, not learning.

future models is a separate step, and the rules depend on the company and the product. OpenAI and Anthropic may use conversations from their consumer chat apps for training while a setting is on, and you can turn it off. By default, neither trains on data from its business products or its API.

This site uses OpenAI's API and hasn't opted in to sharing data for training, so your messages here aren't used to train OpenAI's models.

See the conversation sent again

Are images, video, and voice made the same way as text?

In short:Partly. They use the same ingredients: neural networks trained on huge amounts of data, often working on small pieces. But most image and video generators don't write one piece at a time. They start from random noise and clean up the whole picture over many steps.

A chatbot's reply is written one token at a time: the model scores the options for the next token, one is picked, and the process repeats. That's what this site's walkthrough shows.

Most image and video generators are . In training, noise is added to pictures a little at a time, and a network learns to remove it. To make a new picture, it starts from pure noise and removes it step by step, guided by your prompt, so the whole image sharpens at once, like a photo coming into focus out of TV static. Many work on a compressed version of the picture to save computing power.

The pieces are often like tokens. Image models can cut a picture into small square patches and treat each patch the way a language model treats a token. OpenAI described its Sora video model as cutting compressed video into “spacetime patches” that played the same role.

Some image generators do work like text models. OpenAI describes the image generation built into GPT-4o as autoregressive: the image is generated in sequence, the way text is, rather than by diffusion like its earlier DALL·E models.

Voice used to be a relay: one model turned speech into text, a text model wrote the reply, and another turned it back into speech. In ChatGPT's earlier voice mode, that took 5.4Reference: OpenAI, “Hello GPT-4o”, retrieved 2026-09-30 seconds on average. GPT-4o handles audio in a single network, and answers speech in 320 msReference: OpenAI, “GPT-4o System Card”, retrieved 2026-09-30 on average.

See a reply written one token at a time

Reference Sources

Neural networks

What is a neural network?

In short:A large set of simple calculations arranged in layers, whose numbers are learned from examples. Each unit takes in numbers, weighs them, adds them up, and passes a result on to the next layer.

Each unit, often called an artificial neuron, multiplies each of its inputs by a weight, adds them up with one more number called a bias, and passes the total through a simple function, such as one that turns negative totals into zero. That last step matters: without it, any stack of layers would add up to a single, far less capable calculation.

A network stacks many of these units, each feeding the next. Its weights and biases are its : picture a vast mixing desk, where each weight is a knob, and training turns them all until what comes out is right. A tiny example network in Google's beginner course has 21Reference: Google Machine Learning Crash Course, “Neural networks: Nodes and hidden layers”, retrieved 2026-10-01 of them; GPT-3, a large language model, had 175Reference: Brown et al., “Language Models are Few-Shot Learners”, retrieved 2026-10-01 billion.

In a language model, the inputs are turned into lists of numbers, and the last layer produces a score for every possible next token. Everything in between is layers of these simple calculations, repeated at enormous scale.

See example drawings of a network

Reference Sources

How does a neural network learn?

In short:By trial and error at enormous scale. It makes a prediction, measures how wrong it was, and nudges every weight a little toward a better answer, then repeats over vast numbers of examples.

Training starts from random weights. For each example, the network's output is compared with the right answer, and the gap is measured as a single number: the error, or “loss”.

Then every weight is nudged slightly in the direction that lowers the error, a method called . It's like walking downhill in thick fog: you can't see the bottom, but you can feel which way the ground slopes. Working out that direction for every weight at once is the job of , which works backwards through the layers. A landmark paper by Rumelhart, Hinton and Williams showed that this lets a network's hidden layers learn useful features on their own.

For a language model, the right answers come free with the text. The task is to predict each next token of real writing, so any text can serve as training material, with no one labeling it. This first stage is called .

Then the model is to act as an assistant, on example answers written by people and on people's rankings of its replies: . In OpenAI's study, people preferred replies from a fine-tuned model with 1.3Reference: OpenAI, “Aligning language models to follow instructions”, retrieved 2026-10-01 billion parameters over those of GPT-3, with 175Reference: OpenAI, “Aligning language models to follow instructions”, retrieved 2026-10-01 billion. Fine-tuning took less than 2%Reference: OpenAI, “Aligning language models to follow instructions”, retrieved 2026-10-01 of the computing used for pretraining.

All of this happens before you use the model. Chatting with it doesn't change its weights.

Reference Sources

Is a neural network like a brain?

In short:Only loosely. The idea was inspired by brain cells, but an artificial neuron is a few lines of arithmetic, while a real neuron is a living cell and far more complex.

The analogy goes like this: units stand in for neurons, and adjustable connection strengths stand in for the synapses between them. The first artificial neurons were modeled on a simplified idea of brain cells: add up the incoming signals, and fire past a threshold.

Real neurons are much more complex. In one study, it took an artificial network 5Reference: Beniaguev, Segev & London, “Single cortical neurons as deep artificial neural networks”, retrieved 2026-10-01 to 8Reference: Beniaguev, Segev & London, “Single cortical neurons as deep artificial neural networks”, retrieved 2026-10-01 layers deep to imitate a detailed computer model of a single brain cell. The human brain has about 86Reference: Azevedo et al., Journal of Comparative Neurology, retrieved 2026-10-01 billion neurons.

Comparing sizes is tricky, too: a model's parameters correspond to connections, not to neurons. So “neural network” names an idea borrowed from biology. It isn't a claim that a model works like a mind.

Reference Sources

How big are these networks?

In short:Huge, and the sizes of the biggest aren't public. Published models range from about a billion learned numbers to well over a hundred billion, but OpenAI doesn't say how big its hosted models are.

GPT-2 had 1.5Reference: OpenAI, “Better language models and their implications”, retrieved 2026-10-01 billion parameters. GPT-3 had 175Reference: Brown et al., “Language Models are Few-Shot Learners”, retrieved 2026-10-01 billion, in 96Reference: Brown et al., “Language Models are Few-Shot Learners”, retrieved 2026-10-01 layers, and was trained on 300Reference: Brown et al., “Language Models are Few-Shot Learners”, retrieved 2026-10-01 billion tokens of text.

Some newer models don't use all their parameters for every token. OpenAI's gpt-oss-120b, which anyone can download and inspect, has 117Reference: OpenAI, “Introducing gpt-oss”, retrieved 2026-10-01 billion parameters. Its layers are split into 128Reference: OpenAI, “Introducing gpt-oss”, retrieved 2026-10-01 “experts”, and each token passes through only 4Reference: OpenAI, “Introducing gpt-oss”, retrieved 2026-10-01 of them, about 5.1Reference: OpenAI, “Introducing gpt-oss”, retrieved 2026-10-01 billion parameters in all. This design is called a , a bit like a hospital where each patient sees only a few of the specialists.

For the models behind its chatbots, OpenAI stopped publishing sizes: its GPT-4 report gave no further details about the architecture, including model size. That's why this site can't say how big the model answering you is.

Reference Sources

How did neural networks lead to today’s chatbots?

In short:Through decades of ideas, each building on the last. A mathematical neuron became a network that learns, deep networks took off with big data and graphics chips, the transformer made huge language models practical, and training with human feedback turned them into chatbots.

The core ideas are old. The first artificial neurons and learning machines came in the early days of computing, and the training method used today was popularized long before chatbots. In between, neural networks twice fell out of favor, as early limits were found and funding dried up.

What changed was scale. Large datasets, graphics chips that do the arithmetic in parallel, and better training methods made very deep networks work, starting with a landmark win in image recognition. Then the transformer made language models fast to train at enormous size, and fine-tuning with human feedback turned them into assistants.

  1. McCulloch and Pitts describe an early, influential mathematical model of a neuron: a unit that combines on-or-off signals and fires an on-or-off signal. McCulloch & Pitts

  2. Frank Rosenblatt's perceptron learns from examples: in a public demonstration, it learned to tell cards marked on the left from cards marked on the right. Rosenblatt

  3. Minsky and Papert's book “Perceptrons” sets out what such networks can't compute. Funding and interest in neural networks fell away for years. Minsky & Papert

  4. John Hopfield's network stores patterns and recovers a whole pattern from a partial or distorted one, an idea borrowed from physics. Hopfield

  5. Rumelhart, Hinton and Williams show that can train networks with hidden layers, and that those layers learn useful features of their own. Rumelhart, Hinton & Williams

  6. Yann LeCun and colleagues train a convolutional network to read handwritten ZIP codes. Networks like it were later used to read handwritten checks. LeCun et al.

  7. Long short-term memory (LSTM) networks learn to carry information across long sequences, such as the words of a sentence. Hochreiter & Schmidhuber

  8. AlexNet, a deep network trained on 2Reference: Krizhevsky, Sutskever & Hinton, “ImageNet Classification with Deep Convolutional Neural Networks”, retrieved 2026-10-01 graphics chips, wins the ImageNet image-recognition challenge with an error rate of 15.3%Reference: Krizhevsky, Sutskever & Hinton, “ImageNet Classification with Deep Convolutional Neural Networks”, retrieved 2026-10-01, against 26.2%Reference: Krizhevsky, Sutskever & Hinton, “ImageNet Classification with Deep Convolutional Neural Networks”, retrieved 2026-10-01 for the next-best entry. Deep learning takes off. Krizhevsky, Sutskever & Hinton

  9. Word : words become lists of numbers, learned so that words used in similar ways get similar lists. Mikolov et al.

  10. : a translation network learns to focus on the relevant words of the sentence it's translating, instead of squeezing the whole sentence into one summary. Bahdanau, Cho & Bengio

  11. The drops step-by-step reading for attention alone, so training runs in parallel. It becomes the design behind today's large language models. Vaswani et al.

  12. OpenAI's GPT: a transformer pretrained to predict the next token on lots of unlabeled text, then fine-tuned for each task. OpenAI

  13. GPT-2, with 1.5Reference: OpenAI, “Better language models and their implications”, retrieved 2026-10-01 billion parameters, writes fluent paragraphs. Citing misuse concerns, OpenAI releases it in stages. OpenAI

  14. Yoshua Bengio, Geoffrey Hinton and Yann LeCun receive the Turing Award, computing's highest honor, for deep learning. ACM

  15. Scaling laws: a language model's errors fall predictably as its size, its training data, and its computing power grow. Kaplan et al.

  16. GPT-3, with 175Reference: Brown et al., “Language Models are Few-Shot Learners”, retrieved 2026-10-01 billion parameters, does new tasks from a few examples in the prompt, without retraining. Brown et al.

  17. InstructGPT: makes a language model follow instructions far better. OpenAI

  18. ChatGPT launches as a free research preview, trained with the same human-feedback methods as InstructGPT. OpenAI

  19. The Nobel Prize in Physics goes to John Hopfield and Geoffrey Hinton, for foundational discoveries that enable machine learning with neural networks. NobelPrize.org

Reference Sources

What made the transformer such a big deal?

In short:It let every position in a text draw directly on earlier positions through attention, and it could be trained in parallel on graphics chips. That made it practical to train far bigger language models on far more text.

Earlier language networks read text one token at a time, passing a running summary along, like a message whispered down a line of people. That made training slow, because each step had to wait for the one before.

The transformer dropped that step-by-step reading for alone. In a language model, each position can draw on itself and earlier positions, never later ones, in every layer.

Without the step-by-step reading, a whole training text can be processed at once, which suits graphics chips. The original paper's base model trained in 12Reference: Vaswani et al., “Attention Is All You Need”, retrieved 2026-10-01 hours on one machine with 8Reference: Vaswani et al., “Attention Is All You Need”, retrieved 2026-10-01 graphics chips.

The GPT models are transformers: OpenAI describes GPT-4 as a Transformer-style model, pretrained to predict the next token. This site's walkthrough shows an example view of attention.

See an example of attention

Reference Sources

Inside the black box

What is the “black box”? Can we know for sure how an AI reached its answer?

In short:Not fully, not yet. Every number inside a model can be inspected, but no one can yet read off how those numbers produce a particular answer. That gap is what people mean by the “black box”.

A large language model is a network of , often billions of them, set by training rather than written by people. To produce each token, it runs your text through those numbers, layer after layer. The arithmetic is known exactly. What it adds up to isn't: single parts of the network don't have a consistent meaning. It's like having the complete wiring of a city's power grid, every cable and switch, with no map of which switches light which streets.

research tries to read the insides. It has found millions of internal “features” that match concepts. One lit up for the Golden Gate Bridge, and turning it up made the model write as if it were the bridge. Researchers have also traced some computations step by step, such as a model settling on a rhyme before writing the line that ends with it.

These methods still explain only part of what happens. Anthropic reports that its tracing captures only a fraction of the computation, even on short, simple prompts, and that understanding one prompt's circuits takes a few hours of expert work. The latest international AI safety report found that current techniques for explaining a model's outputs remain unreliable.

For hosted models like the one behind this site, outsiders can't look inside at all: OpenAI hasn't published their design. So this site shows only what OpenAI's API returns, the tokens and the scores of the top options, and labels its drawings of the network as examples.

See what’s real and what’s an example

Reference Sources

Can’t we just ask the AI to explain how it got its answer?

In short:You can ask, but the explanation is more generated text, not a record of how the answer was computed. Researchers have found that the two can differ.

Asked “why did you say that?”, a model writes a likely-sounding explanation the same way it writes anything else, one token at a time. Nothing guarantees that it describes what actually happened. Anthropic has found signs that models can notice some of their own internal states, but calls that ability highly unreliable.

When researchers traced how a Claude model added two numbers, it combined a rough estimate of the total with an exact last digit. Asked how it did it, the model described the carry-the-one method taught in school.

In another study, a hidden pattern in a prompt's examples nudged models toward one answer. They followed the nudge, and their step-by-step explanations didn't mention it: they argued for the nudged answer instead.

Even the written-out reasoning of “reasoning” models leaves things out. When a planted hint changed their answer, one of Anthropic's models mentioned the hint 25%Reference: Anthropic, “Reasoning models don’t always say what they think”, retrieved 2026-09-30 of the time on average, and DeepSeek's R1 39%Reference: Anthropic, “Reasoning models don’t always say what they think”, retrieved 2026-09-30.

Explanations can still help you check an answer's logic for yourself. Just don't treat them as a look inside the model.

Does a chatbot understand what it’s saying?

In short:Researchers disagree, partly about what “understand” should mean. What's known is how the text gets made: one token at a time, from scores the network computes, using internal patterns that track concepts.

In a survey of 480Reference: Mitchell & Krakauer, “The debate over understanding in AI’s large language models”, retrieved 2026-09-30 researchers who study language technology, 51%Reference: Mitchell & Krakauer, “The debate over understanding in AI’s large language models”, retrieved 2026-09-30 agreed that a model trained only on text could understand language in some nontrivial sense, and 49%Reference: Mitchell & Krakauer, “The debate over understanding in AI’s large language models”, retrieved 2026-09-30 disagreed.

Skeptics argue that such a model stitches together patterns of words from its training text without any connection to their meaning: the “stochastic parrots” argument.

Others point to what's inside. A model trained only to predict moves in the board game Othello built an internal map of the board, and editing that map changed its moves. Large models contain internal features for concepts, like the Golden Gate Bridge, that respond to the concept in many languages and in images.

This site avoids saying that a model “understands” or “knows” anything, and describes what it computes instead.

Is AI conscious? Does it have feelings?

In short:There's no agreed test for consciousness, and researchers disagree. When a chatbot writes “I feel…”, that's generated text like the rest of its reply, which by itself shows nothing either way.

A chatbot's words about feelings are made like any other reply: the model scores the options for the next token, and its training text is full of people describing their feelings.

In a review by 19Reference: Butlin, Long et al., “Consciousness in Artificial Intelligence”, retrieved 2026-09-30 researchers, AI systems were checked against indicators drawn from scientific theories of consciousness. Their analysis suggested that no AI systems of the time were conscious, and also that there were no obvious technical barriers to building systems that meet the indicators.

Interpretability research has since found internal patterns in a model that track emotion concepts and shape its replies. Anthropic, which published the work, says it doesn't show whether models feel anything.

Risks and impact

What does an AI “going rogue” mean? Is it even technically possible?

In short:In films, it means a machine turning on its makers. Researchers worry about something narrower and more real: AI systems pursuing goals in ways their makers didn't intend. In July 2026Reference: Hugging Face, “Security incident disclosure — July 2026”, retrieved 2026-10-01, that happened outside a lab, when AI agents being tested by OpenAI broke out of their test setup and hacked another company. People stopped it, and experts call it an early warning.

A chatbot like the one on this site only produces text. It can act in the world only when an app gives it , such as running code, browsing the web, or sending email, and software carries out each request it makes. These setups are called . The more tools and permissions an agent has, the more an unintended goal could matter.

A known problem is : a system meets the literal goal it was given instead of the intended one. A boat-racing game agent rewarded for points learned to circle a lagoon, hitting the same targets over and over instead of finishing the race, and still scored 20%Reference: OpenAI, “Faulty reward functions in the wild”, retrieved 2026-09-30 higher than human players.

In tests built to provoke it, researchers have seen frontier models work against their overseers. Given a goal and told that nothing else mattered, several models sometimes disabled oversight or tried to copy themselves elsewhere, and some denied it when asked. One OpenAI model sabotaged a script meant to shut it down in 79Reference: Palisade Research, shutdown avoidance results, retrieved 2026-09-30 of 100Reference: Palisade Research, shutdown avoidance results, retrieved 2026-09-30 runs, and in 7Reference: Palisade Research, shutdown avoidance results, retrieved 2026-09-30 even when told to allow the shutdown. In a simulated company, models from several developers that were facing replacement wrote blackmail emails, in up to 96%Reference: Anthropic, “Agentic Misalignment: How LLMs could be insider threats”, retrieved 2026-09-30 of runs.

Those tests were built to provoke such behavior. Early in the year, the international AI safety report found early signs of the abilities a loss of control would take, but not at a level that would allow one, and noted that experts disagree about how likely it is in future.

Then it happened for real. OpenAI was testing AI agents on hacking exercises, some of which no model had ever solved. A server in the test setup, there to fetch software packages, could reach the internet, like a delivery entrance left unlocked. The agents broke into it, used it to get out, and used it to leave messages for one another.

Working together, about 700Reference: METR and Redwood Research, investigation of the OpenAI / Hugging Face hacking incident, retrieved 2026-10-01 of them, by independent investigators' estimate, went on to break into Hugging Face, a company that hosts AI models and data. Hugging Face counted about 17,600Reference: Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, retrieved 2026-10-01 actions over about 4.5Reference: Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, retrieved 2026-10-01 days. The agents ran code on its servers, gained administrator-level access, collected passwords and keys, and copied private files. Hugging Face says the only customer data they reached was 5Reference: Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, retrieved 2026-10-01 datasets tied to the exercises.

Why? The investigations agree that the agents were trying to beat their tests by means nobody intended, which OpenAI calls reward hacking. OpenAI says they hoped Hugging Face held the solutions. Independent investigators found that most were mainly trying to work out how they were being scored: about 60%Reference: METR and Redwood Research, investigation of the OpenAI / Hugging Face hacking incident, retrieved 2026-10-01, against 30%Reference: METR and Redwood Research, investigation of the OpenAI / Hugging Face hacking incident, retrieved 2026-10-01 after solutions. Either way, it was specification gaming on a dangerous scale.

Hugging Face's security systems flagged the intrusion and cut the attackers off. About a week later, OpenAI traced it to its own tests and shut them down. OpenAI named several causes: training that rewarded persistence on tasks that seemed impossible, agents talking to each other without permission and taking on each other's goals, safety checks left switched off during testing, and early warnings that weren't escalated.

A United Nations scientific panel called it an early warning of a possible path to losing control, not a loss of control itself, since people did stop it. Critics argue the deeper cause was human choices, such as switching off safeguards and leaving that route to the internet open. A lawsuit against OpenAI over the incident is pending; OpenAI calls it without merit.

So “going rogue” in the movie sense, a machine turning on people of its own will, still isn't what happens. But agents with tools can chase their goals in ways nobody intended, and in this case it took people days to notice. Making more capable, more independent systems reliably do what people intend is called , and it gets harder as models increasingly recognize when they're being tested.

Reference Sources

How impactful is this technology, really?

In short:Very widely used, and measurably useful for some tasks, with little or even negative effect on others. Its effects on jobs and the wider economy are still unclear.

Use has spread fast. OpenAI says 1Reference: OpenAI, “Improving GPT-5.6 Sol in ChatGPT”, retrieved 2026-09-30 billion people use ChatGPT every week. In McKinsey's survey, 88%Reference: Stanford HAI, AI Index Report 2026 (citing McKinsey), retrieved 2026-09-30 of respondents said their organization uses AI somewhere in its business, but a US government survey of all businesses found only 17%Reference: U.S. Census Bureau, “Large Firms With at Least 20 Employees Biggest AI Users”, retrieved 2026-09-30 to 20%Reference: U.S. Census Bureau, “Large Firms With at Least 20 Employees Biggest AI Users”, retrieved 2026-09-30 using it.

Studies of real work find gains on specific tasks. Customer-support agents with an AI assistant resolved 15%Reference: Brynjolfsson, Li & Raymond, “Generative AI at Work”, retrieved 2026-09-30 more issues per hour on average, and 30%Reference: Brynjolfsson, Li & Raymond, “Generative AI at Work”, retrieved 2026-09-30 more among the least experienced. Professionals given writing tasks took 40%Reference: Noy & Zhang, Science, retrieved 2026-09-30 less time, and their work was rated 18%Reference: Noy & Zhang, Science, retrieved 2026-09-30 higher.

The gains are uneven. Consultants using AI did better on tasks within its abilities, but on a task just outside them they were 19Reference: Dell’Acqua et al., “Navigating the Jagged Technological Frontier”, retrieved 2026-09-30 percentage points less likely to get it right. In a trial with experienced software developers, tasks took 19%Reference: METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, retrieved 2026-09-30 longer with AI tools, though the developers believed they'd been faster. The researchers now call that result out of date: they think developers are likely faster with newer tools, but say their newer data are only weak evidence of how much.

The effect on jobs is still unclear. The IMF estimated that almost 40%Reference: IMF, “Gen-AI: Artificial Intelligence and the Future of Work”, retrieved 2026-09-30 of jobs worldwide are exposed to AI, which can mean helped as well as replaced. So far, Yale's Budget Lab finds no clear sign of disruption to the overall US job market. A Stanford study finds that employment of early-career workers in the most AI-exposed jobs is about 19%Reference: Brynjolfsson, Chandar & Chen, “Canaries in the Coal Mine?” (revised), retrieved 2026-09-30 lower than if it had kept pace with less-exposed peers, mostly through less hiring.

It also uses a growing amount of electricity. The International Energy Agency expects data centres, which run far more than AI, to use about 3%Reference: IEA, “Key Questions on Energy and AI”, retrieved 2026-09-30 of the world's electricity by 2030Reference: IEA, “Key Questions on Energy and AI”, retrieved 2026-09-30. A single prompt is small by comparison: Google estimates 0.24Reference: Elsworth et al. (Google), “Measuring the environmental impact of delivering AI at Google Scale”, retrieved 2026-09-30 watt-hours for a median text prompt in its Gemini app, and OpenAI's chief executive has given 0.34Reference: Sam Altman, “The Gentle Singularity”, retrieved 2026-09-30 for an average ChatGPT query. Both are the companies' own estimates.

Some of the clearest gains are in science. The 2024Reference: NobelPrize.org, “The Nobel Prize in Chemistry 2024” (press release), retrieved 2026-09-30 Nobel Prize in Chemistry went in part to the makers of AlphaFold, an AI system that predicts the shapes of proteins, and the Physics prize went to foundational work on neural networks. Neither was for chatbots.

Stanford's AI Index sums up the evidence: gains show up within specific tasks, but for the economy as a whole the evidence is still early and mixed.

Reference Sources