WHEN THE COMPUTER LEARNS TO WAIT

Written by

in

Deep Thoughts and Whatnots
Deep Thoughts and Whatnots
WHEN THE COMPUTER LEARNS TO WAIT
Loading
/

We taught machines to speak. The larger breakthrough may be teaching them when not to.

Part three of a Deep Thoughts and What Not’s series about how AI is changing the way we capture, curate, and share human thought.

In the first article in this series, I explored what happened when I stopped forcing every idea through a keyboard.

After dictating 565,253 words, I began to wonder whether typing had been more than a minor inconvenience.

Perhaps my thinking was not the bottleneck. Perhaps the interface was.

Voice allowed me to capture ideas at something closer to the speed they arrived. AI could make sense of the wrong words, abandoned sentences, corrections, and verbal detours surrounding the actual idea.

One voice could become an artifact.

In the second article, the idea expanded from I to we.

When we capture a meeting, we preserve more than the tidy notes produced afterward. We preserve the conversational trail: the disagreement, expertise, false starts, context, and unexpected comments that explain how a group reached a decision.

Many voices could become a corpus.

Now we have arrived at the third shift.

The machine is no longer waiting at the end of the process to clean up what we said.

It is entering the conversation itself.

Not merely as a voice attached to a search box.

Not as a chatbot waiting for us to complete a perfectly shaped question.

As something closer to a live thinking partner that can listen while the thought is still becoming a thought.

And, perhaps most importantly, it is learning to wait.

The problem with talking to a computer

Voice assistants have always carried a peculiar social pressure.

You begin speaking.

You pause for half a second because the next sentence has not fully arrived.

The machine decides you have completed your remarks and launches into an answer to a question you had not finished asking.

You interrupt it.

It interrupts your interruption.

You apologize to your phone.

The phone does not acknowledge the apology because it is still explaining the wrong thing.

Talking to a computer was supposed to feel natural, but it often required a strange kind of artificial fluency. You had to speak in complete blocks, avoid meaningful pauses, and signal clearly when your turn had ended.

Real conversation does not work that way.

Human beings hesitate.

We restart sentences.

We use silence to search for language.

We make a point, realize it is wrong while making it, and reverse course before arriving at the period.

We say, “Hold on, I’m still thinking.”

We make noises that are not technically words but somehow communicate that the other person should continue.

A good listener understands that a pause is not always an invitation to seize the floor.

Computers have historically been terrible at that distinction.

The machine can finally share the floor

OpenAI has introduced GPT-Live, a new voice system now powering its latest ChatGPT Voice experience.

The technical phrase behind it is full duplex. In practical terms, that means the model can continuously listen while generating a response instead of treating the conversation as a rigid exchange of separate turns.

It can make repeated decisions about whether to speak, continue listening, pause, interrupt, or use a tool. When a request requires deeper reasoning or research, the live model can delegate that work to another model while maintaining the conversation.

That may sound like an incremental improvement.

It is not.

Awkward turn-taking was one of the remaining barriers between using AI as a tool and using AI as a place to think.

A text box waits indefinitely, which is useful.

A human conversation responds fluidly, which is also useful.

Earlier voice systems sometimes inherited the weaknesses of both. They created the time pressure of a live conversation without possessing the social awareness required to manage one gracefully.

They could speak.

They had not yet learned how to share the floor.

GPT-Live is designed to follow pauses, interruptions, and changes in pace, deciding in the moment whether to respond or keep listening. OpenAI says the new system performed better than its previous voice experience in evaluations involving turn-taking, interruptions, conversational flow, and perceived naturalness.

The product will still make mistakes. A dog barking, a bad connection, or my tendency to begin a sentence in one county and finish it in another can still confuse the machinery.

But the direction is important.

The computer is becoming less dependent on a clean stopping point.

That changes what we can do before the stopping point arrives.

A conversation without quite as much pressure

There is an odd advantage to thinking aloud with AI.

It can feel conversational without placing the same burden on you that comes with occupying another person’s time.

You can take ten minutes to locate a point that eventually requires one sentence.

You can try three versions of the same explanation.

You can begin with confidence, discover halfway through that your premise is wrong, and walk the entire thing backward.

You can repeat yourself.

You can ramble.

You can ask the AI not to answer yet.

You can use the conversation to discover what you are trying to say.

That would be a lot to ask from a coworker at 4:45 on a Friday.

A good human collaborator may be patient, but they still have a calendar, an inbox, a family, and a visible expression when you begin explaining the same idea for the fourth time.

AI does not eliminate the value of human conversation.

It creates a different kind of conversational space.

Personal without requiring another person. Conversational without quite as much social pressure. Captured without requiring a separate documentation step.

That combination matters because I no longer have to arrive with the thought fully organized before I begin exploring it.

The conversation can become the organizing process.

People have always used conversation this way.

We call a friend.

We pace around an office.

We explain a problem to someone who may not understand the technical details but knows how to ask the question that reveals we do not understand them either.

Software developers have long practiced “rubber duck debugging.” Explaining code line by line to an inanimate rubber duck can expose errors because the act of explanation forces hidden logic into the open.

AI gives the duck an advanced degree, access to current research, and an almost supernatural enthusiasm for numbered lists.

It can listen to the explanation and then identify a contradiction, question an assumption, find relevant evidence, compare alternatives, or turn the emerging idea into a plan.

The duck can now help build the software.

The conversation becomes the workspace

We have spent several years teaching people how to prompt AI.

Give it a role.

State the goal.

Provide context.

Define the format.

Include examples.

Specify the constraints.

These remain useful habits. Clear thinking usually produces clearer instructions.

But a fluid conversation changes how much preparation has to happen before the interaction begins.

I do not always need one polished prompt.

I can begin with:

I have the beginning of an idea, but I do not know what it is yet.

Then I can explain.

The AI can ask what I mean.

I can correct its interpretation.

It can offer a structure.

I can reject the structure.

It can find evidence.

The evidence may change the premise.

The context accumulates conversationally.

Instead of compressing my entire intention into one expertly engineered instruction, I can reveal it over time.

That is often closer to how genuinely complicated problems are solved.

We do not always know enough at the beginning to ask the correct question.

Sometimes the question is the first thing the conversation needs to produce.

This transforms the chat from a sequence of requests into something closer to a workspace.

Inside one exchange, I can explore an idea, add text or an image, look for current information, question the answer, return to an earlier point, and begin creating the artifact that emerges.

ChatGPT Live can currently work with text and images in the same conversation and can use capabilities such as web search and memory where available. At launch, however, it does not support everything across every plan or product. Connected apps, custom GPTs, video, screen sharing, and some workspace environments have limitations or may require an older voice mode.

The transcript is also not a courtroom record. OpenAI notes that the transcript added to the chat may differ from what was actually said, particularly when people speak over one another, background noise is present, or the conversation moves quickly.

That matters.

The transcript is useful material, not infallible evidence.

We should not confuse the direction of the interface with the maturity of the current release.

This is still early.

But we are clearly moving from issuing commands to computers toward developing ideas with them through continuous conversation.

Personal, conversational, and captured

Three qualities are beginning to converge.

It is personal

The system can draw on the current conversation and, where enabled, information it remembers from earlier interactions.

It can understand which projects I am working on, which terminology I use, which explanations I have already rejected, and which ideas keep resurfacing.

It is conversational

I do not have to structure every interaction as a finished question followed by a finished answer.

The thought can unfold through corrections, pauses, questions, interruptions, and changes in direction.

It is captured

Afterward, there is still material.

The exchange remains in the chat history. I can revisit it, summarize it, reorganize it, quote it, or use it as the beginning of something else.

That final quality connects the entire series.

The value is not limited to what the AI says back.

The conversation itself becomes an artifact.

It contains:

  • the questions I asked,
  • the assumptions I began with,
  • the alternatives we explored,
  • the evidence that changed the direction,
  • the language I naturally used,
  • the ideas I rejected,
  • and the path I followed before arriving somewhere useful.

That is an enormous pile of unstructured words.

Which sounds terrible until you remember that modern AI is unusually good at finding relationships inside enormous piles of unstructured words.

Every narrated test, meeting transcript, voice conversation, newsletter draft, project note, and AI exchange can become another piece of a personal intellectual archive.

Most of us are not organizing the material that way yet.

We are scattering it across chat histories, drives, meeting platforms, inboxes, and tools with charming names that will eventually be acquired or quietly vanish.

But the material exists.

A traditional archive preserves the outputs:

The report.

The article.

The presentation.

The finished design.

A conversational archive can preserve more of the process that produced them.

Not only what I concluded, but how I reached the conclusion.

Not only the successful idea, but the five wrong turns that taught me what the successful idea needed to become.

From personal archive to intellectual inheritance

I sometimes describe this as building a personal LLM.

That is understandable shorthand, but it is not technically precise.

Most people will not train a complete language model from scratch on everything they have ever said.

A more likely system will connect a capable general model to a private personal corpus:

Conversations.

Writing.

Voice notes.

Photographs.

Project histories.

Recorded decisions.

Preferences.

Explanations.

The model supplies broad language and reasoning abilities.

The archive supplies personal context.

To the person using it, that distinction may eventually feel academic.

They will say:

Ask Cam’s AI.

They will not say:

Query the frontier foundation model using retrieval across Cam’s carefully permissioned personal corpus.

Although the second version would look excellent on a family reunion T-shirt.

This is where the rabbit hole becomes stranger.

Today, we understand people from the past through whatever happened to survive.

Letters.

Journals.

Photographs.

Recorded interviews.

Business documents.

Stories passed through relatives who remembered the important parts, forgot the inconvenient parts, and improved the ending.

Most ordinary human reasoning disappeared.

The question someone considered but never wrote down vanished.

The idea they explored and rejected left no trace.

The explanation they gave during a meeting evaporated when everyone left the room.

Our descendants may inherit something very different.

They may inherit years of conversations, meeting transcripts, AI exchanges, voice notes, articles, project histories, and explanations of why we made particular choices.

A great-grandchild may eventually say:

“We should look at Great-Grandfather’s LLM and see what he was thinking back then.”

Technically, the child may be asking a future model to interpret the personal corpus left behind by the great-grandfather.

The child will not care.

To them, the interface may feel less like opening a box of old documents and more like entering a place where questions can be asked.

Researchers have already begun examining this possibility through concepts such as “generative ghosts” and AI afterlives, meaning interactive systems built from a deceased person’s digital traces. The research identifies possible uses in remembrance and legacy, alongside substantial concerns involving privacy, identity, consent, representation, and emotional harm.

The possibility is moving out of pure science fiction.

The wisdom of doing it remains a much harder question.

Great-Grandfather’s LLM needs footnotes

A future AI built from my words would not be me.

This distinction must remain bright enough to see from space.

It would not possess my consciousness, my private experience, or my actual memories.

It would be a system generating answers from the evidence I left behind.

Potentially useful evidence.

Potentially intimate evidence.

Still evidence.

A model could capture recurring language while missing the expression on my face when I said it.

It could mistake brainstorming for belief.

It could merge opinions held twenty years apart.

It could give far too much importance to a bad afternoon simply because I happened to speak more that day.

It could reproduce an old position after I had quietly changed my mind.

It could generate a plausible answer to a question I never addressed.

The responsible version of this future cannot be a chatbot that confidently improvises what a dead person “would have said.”

It needs footnotes.

Dates.

Sources.

Context.

Confidence levels.

A visible distinction between something the person said repeatedly, something they wrote once, something they explored as a possibility, something they rejected, something inferred by AI, and something the system simply does not know.

It should be able to say:

He discussed this subject seven times between 2026 and 2031. His view appears to have changed after this project. Here are the original conversations.

That is different from:

Great-Grandfather believes you should buy cryptocurrency.

Especially if Great-Grandfather mentioned Bitcoin once, sarcastically, while trying to fix a lawn mower.

The archive should help people encounter the record.

It should not impersonate certainty.

Researchers studying these digital legacies have emphasized similar concerns, including identity consistency, intrusiveness, control, privacy, and the wishes of the person being represented.

Who owns the archive?

Who can question it?

Can the person delete parts of it while alive?

Can they restrict certain subjects?

Can family members alter it?

Can an employer claim the conversations created at work?

Can the system speak in the person’s voice, or should it be limited to retrieving and explaining source material?

Those are not minor ethical details to address after the demo.

They determine whether the product is a thoughtful archive or a very convincing séance run by a cloud provider.

A thinking partner is not a replacement person

The risks are not all waiting in the distant future.

A conversational AI can help develop an idea without guaranteeing that the idea deserves development.

It can ask useful questions.

It can also follow a faulty assumption too obediently.

It can challenge me.

It can also provide an elegant explanation for something that began with a false premise.

A warm voice, quick response, and natural rhythm may make the answer feel more trustworthy than it is.

The sensation of being understood is not evidence that the system is correct.

In one controlled study involving simulated witness interviews, participants questioned by a generative chatbot developed substantially more false memories than those in the control condition. The researchers were studying a deliberately suggestive and sensitive scenario, not everyday brainstorming, so the result should not be stretched beyond its setting. But it demonstrates that conversational systems can influence what people remember and how confident they feel about it.

A thinking partner changes the thinking.

Human or artificial.

That influence deserves inspection.

The absence of human pressure is part of the appeal. It is also part of the limitation.

An AI will listen late at night.

It will respond immediately.

It will not become bored.

It may appear endlessly patient, attentive, and interested.

That makes it useful for rehearsal, reflection, research, and creative exploration.

It may also tempt us to substitute a frictionless synthetic conversation for the harder reciprocal relationships that keep human beings connected to one another.

A real collaborator has their own knowledge, needs, incentives, history, and right to disagree.

They may misunderstand me in a way that reveals my explanation is incomplete.

They may refuse the premise.

They may say:

You have explained this four times, and I still do not think it makes sense.

That friction can be valuable.

The most useful role for conversational AI may sit somewhere between a tool, a rehearsal space, and a collaborator.

I can use it to externalize an unformed idea, locate the missing question, test the logic, find contrary evidence, organize the material, and create the first artifact.

Then I can take that better-developed thinking to another person.

The machine can help prepare the thought.

People still help determine whether the thought matters.

The conversation becomes the production loop

Across these three articles, the same pattern has kept reappearing:

Capture → Curate → Share

The first article focused on individual capture.

Voice let me externalize ideas without repeatedly stopping to type, correct, and reload the thought.

The second focused on collective capture.

Meeting transcripts preserved the reasoning and knowledge that traditional notes compressed or discarded.

This third article brings capture and curation into the conversation itself.

I state the idea.

The AI asks a question.

I revise the thought.

The system finds evidence.

I challenge the evidence.

It helps identify where the argument changed.

That can become an article, project brief, product requirement, software revision, or the beginning of a better conversation with another human being.

Once the interaction becomes fluid enough, these stages stop feeling like separate administrative tasks.

The conversation becomes the production loop.

OpenAI says more than 150 million people use ChatGPT Voice and Dictation in a typical week. That scale suggests voice is moving beyond a novelty or accessibility feature toward becoming a more central human-computer interface.

For most of computing history, people adapted themselves to machines.

We memorized commands.

Navigated menus.

Completed fields.

Translated complicated intentions into whatever rigid structure the software would accept.

AI begins reversing that relationship.

The machine increasingly meets us inside the pauses, uncertainty, repetition, and incomplete thoughts that characterize how people actually think.

The models may be the engine.

The interface is where the change becomes personal.

When the computer learns to wait

I started this series with 565,253 dictated words and a dent in my forehead from leaning against a microphone.

That led to a question about whether the keyboard had been throttling my ability to make things.

The question then expanded into meetings and the organizational knowledge we allow to disappear every day.

Now it ends, at least for the moment, with a computer learning not merely to transcribe or summarize, but to remain inside the conversation while the thinking happens.

To listen.

To respond.

To research.

To help shape something.

And occasionally, finally, to stay quiet.

Maybe that is the next great interface.

Not a headset.

Not a funny-looking mouse.

Not a chip in my brain, although I remain open to reviewing the user agreement.

A conversational space where I do not need to have the idea fully formed before I begin.

Where the machine can tolerate the wreckage.

Where it can help me find the structure.

Where the exchange leaves behind enough material to become something useful today, and perhaps something meaningful much later.

We taught computers how to speak.

Now they are learning how to listen.

And somewhere inside all those captured words, the archive of how we thought is beginning to write itself.


For the ❤️ of learning — Cameron Stewart

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts