Monday, September 28, 2026

What Happens When You Hit "Regenerate"

You tap regenerate like it's a cheap retry. The last answer sits there, almost right, and the button looks like an eraser. It isn't. That click does not nudge a sentence. It throws the whole run away and starts over.

This is the part the interface hides. A finished reply feels like an object you could edit. Under the hood it is a path of token-by-token guesses that already happened. Regenerating does not rewind two words. It re-opens the same prompt, with the same model, and predicts every token again from a blank start.

That is why the second answer can change. The model is sampling a path, not retrieving a filed document. A new path can land in a different wording, a different order, even a different claim — and still arrive with the same smooth confidence. Confidence is not a leftover from the first pass. It is produced again, because the system is built to produce an answer, not to remember that you already got one.

The cost is the same story. One click looks tiny. Behind it, the machine lights the network from a cold start. Nodes that were done have to work again. Tokens stream through from the first position, not from the almost-right middle. You do not see the electricity, the hardware, or the repeated work. You see a spinner, then a new paragraph that feels free.

None of this is a glitch. Regenerating is a full recompute by design. If you want a small tweak, you have to ask for a small tweak. If you hit regenerate, you are paying for the entire run a second time — and you are rolling the dice on a new path that can look just as sure of itself as the last one.

So the job of catching it is still yours. The button will not tell you it started from zero. The new answer will not blush. It will just answer.

Did you know that click restarts the entire computation?

Friday, September 25, 2026

Why AI Hallucinates

A 30-second Short can show the gap being filled. It cannot show you how to catch the fill. If you already watched the video, this post is the part that does not fit in a caption: the mechanism, the incentives, a case that actually reached a courtroom, and a checklist you can use the next time an answer sounds too sure.

What's actually happening

A language model is not looking up a fact and then deciding whether to speak. In pretraining it is asked, over and over, to guess the next token in a stretch of text. Grammar, spelling, and common phrasing are easy because they repeat. A one-off detail — a birthday, a case caption, a paper title that appeared once — is not a pattern. When the prompt still demands a completion, the model fills the hole with the closest fluent continuation it has.

That is gap-filling, not lying. There is no inner judge that knows the difference between "this is attested" and "this would sound right." The same machinery that writes a grammatical sentence will invent a plausible citation. The Short's picture is accurate: a small blank, a huge teal network, a confident plug. The blog's extra space is this: the plug is a statistical best guess, not a retrieval from a verified store.

Ask for a name, a number, a holding, or a URL and you are asking it to complete a rare fact. If that fact was weakly represented in training, the completion will still look like the rest of the sentence — calm, specific, and wrong.

What the research says

OpenAI, 5 September 2025 (Why language models hallucinate): hallucinations are plausible but false statements. The authors argue they persist because standard training and evaluation reward guessing over acknowledging uncertainty. Graded only on accuracy, a model that guesses "September 10" for an unknown birthday has a 1-in-365 shot at a point; "I don't know" is a guaranteed zero. Over thousands of items, the guessing model wins the scoreboard.

https://openai.com/index/why-language-models-hallucinate/

The same team (Kalai, Nachum, Vempala, Zhang) posted the paper on 4 September 2025 as arXiv 2509.04664. If a system cannot tell a valid statement from an invalid one, next-word training produces errors the way a classifier produces false positives. Arbitrary low-frequency facts cannot be predicted from pattern alone. Their proposed fix is socio-technical: change how dominant benchmarks score abstention, not only add another hallucination test.

https://arxiv.org/abs/2509.04664

The cost of a fluent gap-fill is not theoretical. In Mata v. Avianca (S.D.N.Y., 2023), a personal-injury brief cited more than half a dozen decisions — Varghese v. China Southern Airlines, Martinez v. Delta Air Lines, and others — that opposing counsel and the judge could not find. ChatGPT had invented them, then insisted they were in Westlaw and LexisNexis. Judge P. Kevin Castel described bogus opinions with bogus quotes. The lawyers were sanctioned.

https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-lawsuit-chatgpt.html

A confident citation is exactly the plug. The courtroom is what happens when nobody checks the hole.

Why it can't just be "fixed"

People treat hallucination as a bug to patch: more data, a bigger model, a retrieval plugin, a "don't make things up" system prompt. Those help at the margin. They do not remove the incentive.

Pretraining never labels each sentence true or false. It only asks for fluent continuation. Then evaluation, the thing labs actually optimize for, mostly scores exact-match accuracy. A model that abstains looks worse than a model that guesses, even when the guesser is wrong more often in the cases that matter. Leaderboards that give zero for silence and one for a lucky hit will keep producing systems that sound sure.

You can fine-tune a model to say "I'm not sure" on some prompts. You cannot make "I'm not sure" win an accuracy-only exam. Until the tests that define "better" stop punishing honesty, the product that ships will still fill the gap. OpenAI's paper is explicit: this is not solved by one more hallucination benchmark on top of the old scoreboards.

How to spot and handle it

Treat fluent specifics as unpaid invoices. Before you reuse an AI answer:

  • Names, dates, dollar amounts, case captions, paper titles, and URLs get checked in a primary source — a docket, a journal page, a company filing — not by asking the same chat "are you sure?"
  • If the model cites a case or a paper, search the identifier yourself. Mata is what happens when verification is another prompt to the inventor.
  • Ask for the source, then open the source. A quote that cannot be found in the linked page is a fill.
  • For anything you would sign, file, or publish, assume the gap was filled until a human-readable record says otherwise.

None of that makes the model honest. It makes you the person who still owns the merge.

Closing

The next generation of models will be smoother. The scoreboards that train them may still pay for a guess. The useful question is not whether the gap will keep getting filled — it will — but whether your workflow still has a human who is allowed to leave the answer blank.

Thursday, September 24, 2026

AI Memory: Will Chatbots Remember Everything?

The old chatbot forgot you every time.

This one already knows your name, your job, last week's trip.

You open a new chat.

You never said remember this.

In the background, it rewrote what it knows about you.

That's persistent memory.

Helpful when it's right.

A problem when it's wrong, or stale, or you never asked it to keep that.

You can still open it, edit it, or turn it off.

An assistant that remembers everything only works if you stay in charge.

What changed

Chatbots used to treat every new thread as a blank slate. Persistent memory is the opposite: the assistant carries a running picture of you into the next chat.

OpenAI, 4 June 2026 (Dreaming: Better memory for a more helpful ChatGPT): ChatGPT's memory system now synthesizes a reviewable profile in the background, instead of waiting for you to say "remember this." The company said the update started with Plus and Pro users in the US, with Free and Go and more countries to follow.

https://openai.com/index/chatgpt-memory-dreaming/

OpenAI's Memory FAQ: you can review, change, or remove remembered information, use Temporary Chat, or turn memory off. Where the option is available, you can also switch back to the older saved-memories list.

https://help.openai.com/en/articles/8590148-memory-in-chatgpt

The part people miss

Remembering is not the same as asking.

A background process can keep preferences current. It can also keep a stale trip, a wrong job title, or a detail you never meant to store. The product only works if the memory is visible and editable.

Stay in charge

The old chatbot forgot you every time.

This one already knows you.

You can still open the memory, edit it, or turn it off.

Wednesday, September 23, 2026

Why AI Needs So Much Computing Power

You typed one sentence.

But AI didn't just read it.

It turned your words into numbers.

Then millions of calculations started happening across powerful chips.

And this happens again and again, layer after layer.

That's why AI needs so much computing power.

And training the model takes even more.

The surprising part?

You see the answer in seconds.

But behind that answer, there's a lot of computation.

What happens behind the scene

AI does not look up a stored reply. It computes one.

The International Energy Agency, 16 April 2026 (Key Questions on Energy and AI): data-centre electricity grew 17% in 2025 to about 485 terawatt-hours. AI-focused sites grew 50% in that year. The IEA's central path has data-centre electricity roughly doubling to about 950 TWh by 2030 — around 3% of global electricity. AI-focused sites triple in the same window.

https://www.iea.org/news/data-centre-electricity-use-surged-in-2025-even-with-tightening-bottlenecks-driving-a-scramble-for-solutions

A simple text query is still small. An agent that reasons is not. IEA 2026 estimates cited by Our World in Data: a standard AI-agent request is about 1.1 watt-hours. An agentic request with reasoning is about 50.

https://ourworldindata.org/how-much-energy-do-data-centers-and-artificial-intelligence-use

The bill that hides in the answer

Gartner, 17 August 2026: inference cost per agentic workflow will rise more than fivefold through 2028. Routing a task to an agentic reasoning model costs a provider at least five times a basic chatbot interaction — often more as the task grows.

https://www.gartner.com/en/newsroom/press-releases/2026-08-17-gartner-predicts-ai-inference-costs-per-agentic-workflow-will-increase-more-than-fivefold-through-2028

Token prices fall. The work per question grows. That is why a tiny question can still need a surprisingly large amount of compute.

Ask, then see the work

You typed one sentence.

You see the answer in seconds.

Behind that answer, there's a lot of computation.

Tuesday, September 22, 2026

AI Agent Did the Job While the Human Watched

I gave an AI agent one simple task: update this report.

It opened the spreadsheet.

Checked the numbers.

Updated the report.

And sent it for review.

I didn't do any of it.

That's the difference between a chatbot and an AI agent.

A chatbot gives you an answer.

An agent actually does the work.

What changed

A chatbot stops at the reply. An agent finishes the ticket.

Gartner, 15 April 2026: agentic AI is at the Peak of Inflated Expectations. Only 17 percent of organizations have actually deployed AI agents. More than 60 percent say they will in the next two years.

The same firm says 40 percent of enterprise applications will have task-specific agents by the end of 2026, up from under 5 percent in 2025. Of the thousands of vendors now calling a product an "AI agent," only about 130 are verifiably agentic. The rest is agent-washing: a chatbot with a new label.

The part people miss

Watching is not the same as owning the outcome.

Gartner's February–March 2026 customer survey (3,566 B2B and B2C): among people who already use generative AI, 58 percent had used it to complete a task on their behalf. In B2B, 74 percent. They want the booking, the submit, the account update — not another FAQ.

Stack Overflow, late April 2026: 63 percent of technologists still rarely or never let an agent run entirely on autopilot.

So yes: the agent can do the job.

The human still has to decide what "complete" meant.

Ask, then watch

I gave an AI agent one simple task: update this report.

I didn't do any of it.

A chatbot gives you an answer.

An agent actually does the work.

Monday, September 21, 2026

The Agent Can Write the PR

AI wrote the code.

It compiled.

All green.

Then we noticed something was missing.

The authentication check was gone.

That's the problem with AI-written code. The agent can write the PR. You still own the merge.

What changed

An agent can close a ticket. A green check is cheap. A merge is still a decision.

JetBrains asked more than 15,000 professional developers, May through July 2026, how last month's work was written. On average they said about 47 percent of the code was fully agent-written. One in five writes none of it without AI. Nine in ten used a coding agent at work at least weekly.

That is not the same as handing the repo to autopilot. Stack Overflow's April 2026 pulse: 63 percent still rarely or never let an agent run entirely on autopilot.

The part people miss

The typing got cheaper. The review did not.

You still need a spec.

A test.

Someone who notices the auth check is gone.

And someone who owns the blast radius when it ships.

Diff, then merge

So yes: the agent can write the PR.

But you still own the merge.

Friday, September 18, 2026

AI Can Generate the Clip

A generated clip beside a film crew

It hasn't replaced the production set.

Today a prompt can return a shot that used to need cameras, locations, actors, and a visual-effects ticket. That is real. It is also the part people over-read.

A clip can look finished. A film is still a pile of decisions.

What changed

AI can create realistic shots. Platforms noticed. Some now label photorealistic generated video even if the uploader does not. Others put limits on wholly generated clips in the feed they pay. Companies are already using the tools in advertising, entertainment, and post-production — the cheap cut, the crowd that would have been left out, the establishing plate.

None of that is the same as handing the day to a prompt.

The part people miss

Generating one impressive eight-second clip is not the same as making an eight-minute film.

You still need a story.

Continuity.

Direction.

Editing.

Sound.

Performance.

And thousands of decisions between the first idea and the final cut. An eight-second loop does not carry those. An eight-minute short still has to.

Clip, then crew

So yes: AI can generate the video.

A clip is not a crew. A prompt is not a production. The set is still a set — camera, lights, someone who owns the continuity, someone who says cut.

That is where the real story of AI-generated video begins. Not at the thumbnail. At the work the thumbnail hides.

What Happens When You Hit "Regenerate"

You tap regenerate like it's a cheap retry. The last answer sits there, almost right, and the button looks like an eraser. It isn't....