Why do I lose the thread when someone talks for more than two minutes?

You lose the thread because live speech is transient. It is gone the moment it is said, so each new point has to be held in working memory until the next one lands, and that store holds only a few items. A document waits for you to come back. A person does not.

Two different failures live inside that complaint, and they have different fixes. The first is overflow: working memory holds only about four chunks at a time, so a dense explanation can fill it before the speaker reaches the point. The second is transience: speech vanishes as it is made. Overflow is about how much you can hold. Transience is about whether you can go back at all. You can work around overflow by holding less. You cannot fix transience by trying harder, because the information you missed no longer exists to be checked.

Cognitive scientists call this the transient information effect. In a 2012 experiment in Learning and Instruction, Wong, Leahy, Marcus and Sweller compared the same content delivered as text and as a spoken or animated sequence. Short segments behave fine. Once the material runs long, the spoken version stops helping and starts hurting, because working memory cannot hold a large amount of information that keeps disappearing. Transient information is anything that was never recorded so you could return to it, which is why writing and recorders exist.

That is why a document feels easier than a person. A 2022 meta-analysis of 46 studies found that reading beat listening only when readers controlled their own pace, and the advantage disappeared when the experimenter set the speed. Pacing is where reading gets its structural edge. When you read, you set the rate. When someone talks, they do.

Why can I follow a video but not a live explanation?

Because a recording is permanent information and a person is not. A video or podcast stays where it is and replays on demand. Live speech exists only in the seconds after it is spoken. You are not better at audio than at people. You are better at anything you can replay, and the repair costs nothing.

Echoic memory, the brief high-fidelity trace of what you just heard, lasts only a few seconds. Research on echoic memory puts the lifetime at a few seconds, and Darwin, Turvey and Crowder first described the effect in 1972. "What did you say?" works inside that window. A question at minute twenty cannot repair minute two, because by then you are asking the speaker to rebuild the explanation from scratch, not replay a recording in your head.

Text also lets you regress invisibly. A re-read of a confusing sentence is free and nobody sees it. A re-ask of a person is neither. Speech makes every regression visible, which is why the repair gets expensive and why people quietly stop making it.

How do I track a long explanation without interrupting?

Stop trying to hold every point. Capture one headline: the single sentence that says what the speaker is driving at so far. Then listen for how each new point connects back to it. You are tracking one thread, not a transcript, and asking one targeted question when the link breaks costs far less than a full restart.

The social cost is real, and most advice ignores it. An r/AutisticAdults thread with 107 upvotes puts the whole pattern in one sentence: "Somebody says something and it takes you like five seconds to figure out wtf they just said. I feel like I nod and smile instead of hearing correctly way too often... But people also get annoyed when you ask them to repeat themselves." The nod is not laziness. It is a workaround, priced correctly, for a repair that costs status every time you make it.

Some of what you experience as your own failure is a delivery problem wearing your name. A 473-upvote r/managers thread describes a team member who explains every small thing in depth until meetings have to be cut short. An r/PetPeeves post prices the delay exactly: two minutes of signal after eight minutes of fluff. When the thread breaks in a meeting like that, nothing about you changed. The format did. That distinction matters, because it tells you whether to work on your listening or on getting the information in a form you can re-scan.

The practical move is to stop trying to hold the whole explanation and hold the link instead. Write or type the one headline of what the speaker has said so far, then listen for how the next part connects to it. If the thread still breaks, spend your one interruption on the step you lost, not on the whole talk. "Can you say that last bit again?" is a request. "Can you start over?" is a cost you both pay. One is a repair. The other is a re-do.

That habit, capturing the headline and listening for the link, is trainable in a way most listening advice is not, because it is a behavior you can practice and measure. Absorb is a short daily workout built for this: a few minutes of listening, reading and visual challenges that show you what you actually took in, with scores that stay on your device. You find out which channel drops first when the material gets dense. see Absorb on the App Store