How we know
Sound gets to your inner ear by two physical paths. The first is air conduction. Your vocal folds vibrate molecules of air, those waves travel out of your mouth, bend around your head, and arrive at your eardrum. That is the only path a microphone captures. The second is bone conduction. While you are speaking, your vocal folds also vibrate the bones of your skull and jaw directly. Those vibrations reach the cochlea through the bone itself, bypassing the eardrum entirely.
The two paths are not identical. Bone conduction is dramatically better at carrying low frequencies than air conduction is. That is the bass you feel when you hum with your fingers in your ears. So when you speak out loud, the cochlea receives two simultaneous signals. A thin, high-frequency-skewed air signal, and a thick, bass-rich bone signal mixed underneath. Your brain combines them into the voice you have heard inside your head for as long as you can remember.
In 2023, Katarzyna Pisanski and colleagues at the Royal Society Open Science put the mechanism to the cleanest test yet. They recorded the speech of each participant and played those recordings back two ways. Through normal headphones, which deliver air conduction only. And through a bone-conduction transducer placed against the skull, which restores the missing low-frequency component. Participants were asked which clip sounded most like their own voice. The bone-conduction-enriched version won almost every time. Strip the bone path out and your voice does not feel like your voice.
The physics behind both paths was mapped by Georg von Bekesy in the work that earned him the 1961 Nobel Prize in Physiology or Medicine. Modern hearing aids and bone-conduction headphones rely on the same principles.
What it means
Two practical takeaways. First, the version on the recording is not a distortion. It is the air-only version of your voice, which is exactly what every other person has been hearing since the day you met them. Second, the version inside your head is also real, just bone-enriched. Neither is a fake. Your brain has simply been listening to a richer mix than your friends and family have.
You can demonstrate it in five seconds. Cover both ears tightly and hum. The hum you hear is almost entirely bone conduction, with the air path muted. Now uncover your ears and hum at the same volume. The bass thins because the air path is back in the mix but it does not match the bone-rich version. The first sound, with ears covered, is closer to the voice in your head. The second is closer to your recorded voice.
Why it feels worse than it is
There is also a familiarity gap. You have heard your bone-rich voice every day of your life. You have rarely heard your air-only voice. The first time you hear a recording, your brain runs a familiarity check, fails, and flags the sound as wrong. The voice is not actually worse. It is just unfamiliar. People who do voice work for a living stop noticing the gap after a few months of regular playback.