Before any human could read a single word, every human could tell a story, and anyone within earshot could receive it. That asymmetry is close to 200,000 years old on one side and roughly 5,000 years old on the other. Spoken language, as far as the fossil and genetic evidence can tell us, is old enough to be load-bearing in our neurology. Writing is a tool we bolted on afterward: a clever hack that happens to work, not a native capacity we were born wired for.
I have spent the better part of three decades building software at the intersection of biology and computing, so I am not anti-text. I am typing this sentence, after all. But I think we have quietly confused the accelerant for the engine. The printing press, the paperback, the inbox, the feed: all of them move information faster than a human voice ever could across distance and time. None of them changed what kind of animal we are. We are, underneath the literacy, still a species built to speak and to listen. The current wave of voice technology, dictation that replaces typing, text-to-speech that reads to us in our own cloned voice, real-time translation that erases language barriers mid-sentence, looks to me less like a novelty and more like a return trip.
Speech Is Older Than the Species Doing the Speaking
Modern Homo sapiens is roughly 300,000 years old. The best estimates for the emergence of complex spoken language, built on genes like FOXP2 The evolutionary history of genes involved in spoken and written ... and a vocal tract capable of the fine articulatory control speech requires, put the capacity somewhere between 100,000 and 200,000 years ago. Writing, by contrast, is about 5,000 years old, invented independently a handful of times, first in Sumer for keeping track of grain and beer.

That gap matters because of how each capacity is acquired. A child raised in a normal linguistic environment acquires spoken language without being taught, on a predictable developmental timetable, using dedicated neural architecture (Broca’s and Wernicke’s areas) that specialized for exactly this job over evolutionary time. Reading has no dedicated brain region. The neuroscientist Stanislas Dehaene has described literacy as a case of ‘neuronal recycling’ Inside the Letterbox: How Literacy Transforms the Human Brain - PMC, in which circuitry built for object and face recognition gets repurposed, late in each individual’s development and only with years of explicit instruction, to decode strings of symbols. That is why learning to read takes a classroom and a curriculum, and learning to talk takes a crib and a few thousand hours of ambient exposure.
Dyslexia is, in a sense, the tell. It is not a defect in a dedicated reading organ, because there isn’t one to be defective. It is friction in the retrofit.
The Printing Press Was an Accelerant, Not a New Engine
None of this is an argument against writing. Gutenberg’s press, somewhere around 1440, did something genuinely extraordinary: it took knowledge that used to travel at the speed of a human voice, one copied manuscript at a time, and let it propagate at the speed of a printing run. Literacy rates climbed. The Reformation happened. Modern science, which depends on claims being checked, cited, and reproduced across distance and time, becomes possible at a scale that oral transmission alone could never support.
But notice what the press actually did. It did not change the brain doing the reading. It built a faster delivery mechanism on top of the same repurposed visual circuitry, and it asked more of that circuitry, earlier in life, more universally, than had ever been asked before. The printing press is fire: an external technology that extends a biological capacity (vision, in this case) far past its native range, without ever becoming part of our biology. We are now layering large language models and AI writing assistants on top of that same stack, accelerating text production and editing yet again. Each layer increases the speed of knowledge transfer. None of them touches the underlying wiring, which is still built for sound.
Voice as the Great Leveler
Here is the practical consequence of treating speech as the default and literacy as the add-on: almost everyone, barring specific pathology, acquires spoken language. Global literacy, even after a century of universal schooling as a stated policy goal, still sits in the high 80s as a percentage Literacy - UNESCO, and that number drops sharply in regions with limited access to education, and historically has always been lower for women and the poor than for the groups who got to decide what counted as literate. Literacy has always been a gate, and gates have always had gatekeepers.
Voice does not require a gate in the same way. A toddler in Lagos and a toddler in Oslo both learn to talk before either learns the alphabet of their own language, let alone anyone else’s. Modern voice technology is starting to exploit that shared substrate: real-time speech translation that lets two people in different languages have an actual conversation, dictation tools that let someone who struggles with a keyboard or a spelling system get their ideas into a document anyway, text-to-speech that opens a book to someone who cannot or does not read print fluently. None of these tools are workarounds for a deficiency. They are routing information through the channel we were actually built to use, instead of insisting everyone first learn the 5,000-year-old hack.
Why Reading Right Before Bed Might Be Fighting You
Here is the part I expect some pushback on. During REM sleep, the stage most closely associated with dreaming and with consolidating emotional and declarative memory The REM Sleep - Memory Consolidation Hypothesis - PMC - NIH, your eyes move rapidly and more or less at random: no fixed target, no task, no line to track. Sleep researchers have spent decades trying to pin down exactly what those movements are doing, but the correlation between REM-stage eye activity and memory consolidation is well established in the sleep science literature.
Separately, and for different reasons, therapists using EMDR (eye movement desensitization and reprocessing) have people move their eyes left and right, sometimes up and down or in slow circles, while they process difficult memories, reporting that the bilateral movement itself seems to lower arousal and ease distress. People have independently stumbled onto versions of this as a pre-sleep ritual: deliberate slow eye movements, left-right, up-down, counter-clockwise, clockwise, meant to settle the nervous system before the lights go off.
Now put reading next to that. Reading is the opposite of free eye movement. Your eyes move in short, fast jumps called saccades, each one landing on a word for roughly a quarter of a second, almost entirely left to right if you read English, backtracking only to fix comprehension errors. It is a tightly controlled, goal-directed, linear ocular task: the mechanical inverse of the loose, undirected movement associated with REM and with the calming techniques above. I don’t have a controlled study that proves reading-before-bed undermines rest (if you know of one, send it my way), but the directionality is suggestive enough that I have started treating ‘read a book to wind down’ as a less obviously correct piece of advice than it’s usually presented as. Listening to a story, by contrast, asks nothing of your eyes at all. You can close them.

Typing Is a Very Small Movement for a Very Big Brain
There is good reason to believe that fine motor dexterity, the manipulation of objects, tool use, the fingers doing something skilled and varied, played a real role in the expansion of the brain regions we now use for planning and language Research progress on the relationship between fine motor skills and .... Hand-brain co-evolution is a serious line of research, not a folk theory: the same cortical real estate that handles complex finger movement sits immediately next to, and appears to have scaffolded, language production areas.
Which makes it a little strange that our main daily interface with information now asks almost nothing of the body. Typing is dexterous, technically, but it is also narrow: the same small, repeated finger movements, hour after hour, torso still, eyes locked on a screen twenty inches away. Compare that to what an oral storyteller’s body is doing: standing or pacing, gesturing with both arms, modulating pitch and pace and volume, making eye contact, responding to a listener’s face in real time. Walking itself has a measurable effect on idea generation. A frequently cited Stanford study found that walking, compared to sitting, reliably boosted creative output on divergent-thinking tasks, with the effect persisting even after the person sat back down Stanford study finds walking improves creativity. Oral storytelling was never a seated, finger-only activity. It used a body. I think there’s a real argument that our current default (sit still, move only your fingers) is a narrower channel than the one we evolved to think and communicate through, and that speaking while moving, pacing while dictating, is closer to the old default than typing at a desk will ever be.
Top 10 Toy Stories: Oral Tradition in Miniature
If you want to see the oral tradition still operating in the wild, you do not need to go to an archive. You need to watch how stories actually get told when nobody is writing them down. Here are ten forms, old and new, where the story lives in the voice and the gesture, not on the page.

Aesop’s Fables, told aloud. Centuries before anyone wrote them down, these animal morality tales were spoken entertainment, compressed, punchy, built to survive being repeated from memory.
Wayang kulit shadow puppetry (Indonesia).A single puppeteer narrates, voices every character, and cues a live gamelan orchestra, for hours, from memory, behind a lit screen.
The West African griot tradition. A griot carries a family’s or a kingdom’s entire oral history, often set to the kora, functioning as historian, genealogist, and entertainer in one voice.
The Homeric epics, before they were epics on a page. The Iliad and the Odyssey existed as sung, memorized performance for generations before being fixed in writing, using meter and repetition specifically because those features help a human memorize and perform a 12,000-line poem.
Finger puppets and bedtime stories. The smallest, most universal version: a sock or a felt puppet, a parent’s voice doing three characters, no text required.
Playground chants and counting-out rhymes.Passed child to child with zero adult curation or written transmission, mutating slightly in every schoolyard, a living oral-tradition laboratory.
The talking-stick storytelling circle. Used in various Indigenous North American traditions, where whoever holds the stick holds the floor, and the story is shaped by the room in real time.
Jewish oral Torah and midrash tradition.Interpretation and commentary were transmitted by recitation and memorization for centuries before the Mishnah was finally written down.
The campfire ghost story. No author, no fixed text, endlessly reworked by whoever is holding the flashlight under their chin that year.
The voice-cloned bedtime story app. A parent records (or clones) their voice once, and a child can summon a new story in that exact voice on demand, even when the parent is traveling. The oldest form of storytelling, closing the loop through the newest technology.
The Toolkit for Going Back
If the thesis is right, that our native mode is speaking and listening and reading/writing is the add-on, then the interesting question is what tools let us lean back into the native mode without giving up what literacy built. I’ve been using a few of these daily.
Rambleproof (rambleproof.com) is built around the idea that you should be able to just talk, including rambling, circling back, restarting a sentence halfway through, and still end up with clean, usable output. That matters because it matches how speech actually works: unedited, associative, looping back on itself, nothing like the linear sentence you’d write. Most dictation tools assume you’ll speak the way you’d type. Rambleproof assumes you’ll speak the way you actually speak, then does the cleanup. Rambleproof also critically runs locally on your machine without having to call out to the cloud. I think this is critically important in the age of privacy. You never know what those cloud-based providers are doing with what you say, and nowadays, since I speak more than I type, I’m especially sensitive to this. Rambleproof is now my go-to.

Wispr Flow (wisprflow.ai) was the initial technology that I used after hearing it on my favorite podcast, My First Million. I noticed it struggled when I was having internet connections, but when we changed to fiber, that became less of an issue. It always nagged me that it was going out to the cloud. You never know with these startup business models. Wispr Flow knows a ton about me, more than any social media or Google search engine, because I used it in my day-to-day on my desktop and on my iPhone. it’s also extremely difficult and laborious to download the audio clips that they have of me. What are they doing with it, anyway. Rambleproof is a great example of bringing things down locally, running models on far less powerful machines that are running in the cloud. This is a big tell. This is where AI is headed: bringing powerful inference models locally. It’s another reason why we’re not going to need to build all those data centers, and especially not need to build data centers in space, but that’s a whole other article for next time.
Then there’s the other half of the loop: hearing it back. ElevenLabs and similar voice-cloning tools can take a short sample of your own voice and use it to read anything back to you, your own draft, someone else’s article, a book you’d rather listen to than see. When you read silently, you almost certainly hear an internal voice doing the reading, the ‘inner voice’ that linguists and psychologists have documented as a near-universal feature of literate reading. I’d argue that voice, generated by a model trained on your actual speech, carries more weight than the flattened inner voice of silent reading, partly because of the well-documented self-reference effect in memory research Memory for Details with Self-Referencing - PMC - NIH, where information processed in relation to yourself gets encoded more deeply than information processed in the abstract. Hearing your own words, in your own voice, read back to you, is about as self-referential as input gets. My own habit now: dictate the rough draft with Rambleproof, tighten it in text, then listen to the final pass in my cloned voice before I publish. I catch different errors that way than I catch by rereading, and the content sticks with me longer afterward.
My Voice Is Also a Lab Sample
There is a second reason I keep my recordings local, and it goes beyond privacy. Every Rambleproof session leaves the raw audio on my own disk. Over months that adds up to hours of one speaker on one microphone, recorded morning and night, fed and fasted. That is a longitudinal voice dataset on a single person, and I own every second of it. It also answers the question I asked about Wispr Flow a few paragraphs ago. One thing you can do with a large archive of someone’s voice is read their health from it.
That sounds like a stretch until you look at what voice researchers actually measure. Your vocal folds vibrate a hundred or more times a second, and no two cycles are identical. Jitter is the cycle-to-cycle wobble in pitch. Shimmer is the cycle-to-cycle wobble in loudness. Harmonics-to-noise ratio (HNR) measures how much of the sound is clean tone versus breathy noise, and cepstral peak prominence (CPPS) captures how clear and periodic the voice is overall. You would never hear these consciously. The larynx is muscle, nerve, and blood vessel like the rest of the body, so systemic disease can reach it, and these numbers move when it does. Parkinson’s research has relied on them for years, because speech changes are among the earliest measurable signs of the disease.

Diabetes is the newer frontier. A conversation with a company (Rhema Health) working in this space sent me into the literature, and the signal is there:
Gestational diabetes and shimmer. Among 101 pregnant women tested at 24 to 28 weeks, the 31 with gestational diabetes had significantly higher shimmer and lower HNR and CPPS than controls, while pitch and jitter were unchanged. Shimmer alone separated the groups with an AUC of 0.72 (Hepkarsi et al., Journal of Voice, 2026).
Hypoglycemia from a smartphone. A Bern group recorded 540 smartphone voice samples from 22 adults with type 1 diabetes under controlled normal and low blood sugar. A machine learning model detected hypoglycemia from voice alone with an AUC of 0.90 when people read a passage aloud (Lehmann et al., Diabetes Care, 2026).
Type 2 diabetes from everyday sentences. In 461 adults reading random sentences on their own phones, a spectrogram transformer classified type 2 diabetes status from voice alone with an AUC of 0.78 for men and 0.79 for women (Guermazi et al., IEEE JBHI, 2026).
The older literature explains why the newer results look the way they do. A 2024 meta-analysis pooling five case-control studies, 321 people with type 2 diabetes and 171 controls, found no significant group difference in average pitch, jitter, shimmer, or HNR (Hamdan et al., Folia Phoniatrica et Logopaedica, 2024). One measurement per person, averaged across strangers, washes the signal out. It appears when a model sees many recordings per person, and the strongest result above is the one that compares the same people to themselves at two different blood sugar levels.
What I have not found anywhere is the longitudinal version: one person’s voice tracked over months against their own HbA1c and continuous glucose readings. Every study above is a snapshot. That is the gap my own recordings can start to fill. I already pair continuous glucose data with my meal logs in Food Health TxD, and Rambleproof is quietly building the matching voice record on the same timeline. The plan is to extract shimmer, jitter, HNR, CPPS, and spectral features from each session, line them up against running glucose windows and HbA1c, and see what tracks. A random forest on those hand-built features comes first, then a small neural network once there are enough labeled sessions. A model like that is small enough to ship inside TxD and run on the phone, so anyone could dictate their meal notes, build a voice-and-glucose training set from their own data, and get a personal model back without a single recording leaving the device.

This is the whole article in miniature. The oldest channel we have, the human voice, carries information about the body that a keyboard never captured. Keeping it local is what makes it mine to learn from.
A Return Trip, Not a Retreat
None of this is an argument for abandoning the written word. I am not going to stop reading, and you should not either. But I think we have spent five centuries treating literacy as the upgrade and orality as the primitive version, when the honest evolutionary story runs the other way. Speech and listening are the old, deep, universal operating system. Reading and writing are a brilliant, genuinely world-changing piece of software we installed on top of it, one that still requires years of training to run and that a meaningful share of humanity never fully gets access to.
Voice technology is, for the first time, cheap and good enough to route information through the older channel at the speed the newer one made possible: dictation instead of typing, translation instead of a shared alphabet, your own cloned voice instead of a flattened inner monologue, a story told instead of a page turned. That is not nostalgia. It’s catching the delivery mechanism up to the biology it was always supposed to serve.
References
REM sleep and memory consolidation: The REM Sleep - Memory Consolidation Hypothesis - PMC - NIH
Reading as ‘neuronal recycling’ (Dehaene): Inside the Letterbox: How Literacy Transforms the Human Brain - PMC
Hand dexterity and brain development: Research progress on the relationship between fine motor skills and ...
Walking and creative thinking (Oppezzo & Schwartz): Stanford study finds walking improves creativity
Self-reference effect in memory encoding: Memory for Details with Self-Referencing - PMC - NIH
Global literacy rate data (UNESCO): Literacy - UNESCO
FOXP2 and the emergence of spoken language: The evolutionary history of genes involved in spoken and written ...
Ramble Proof: https://rambleproof.com/
Wispr Flow: https://wisprflow.ai/
ElevenLabs: https://elevenlabs.io/
Food Health and Food Health TxD: foodhealthscan.com
Steven Muskal, Ph.D. is the CEO of Eidogen-Sertanty, Inc. - a drug discovery informatics company. He has spent four decades working at the intersection of computational biology, AI, and drug discovery. He writes about AI, health, and the intersection of biology and technology at stevenmuskal.com
For a clip from a recent mix with Maya, Tim, David, Don, and Dom. This one is one of my favorites - David started a bass line, Tim suggested Roller Coaster, and Maya took charge and crushed Just a girl. Priceless!




