How to turn old voice recordings into a written memoir
The path from a box of cassettes, voicemails, and phone recordings to a printed book: digitizing, verbatim transcription, and chapters built from fragments.
By The Yourtale team · Published 8 September 2026 · 13 min read
Turning old voice recordings into a memoir runs in a fixed order. Digitize the audio first, because the tape is the part with a deadline. Transcribe it verbatim rather than cleaned up. Index every recording by what it actually covers, then build chapters around the material you have instead of the life you wish had been recorded. Edit against the audio, mark the gaps honestly, and print the book with the recordings kept alongside it.
Most guides to this stop at "use a transcription service." That is one step out of six, and it is not the hard one. The hard one is that these recordings were almost never interviews. They are a birthday dinner, a car ride, a voicemail about a hospital appointment, forty minutes of your grandfather talking about the war because someone happened to leave a recorder running. They are fragments, in no order, about whatever came up. And the person who could fill the gaps is usually gone, which is why you are doing this at all.
That single constraint changes the whole method. This piece is the method.
We build a memoir service, so we have a stake in the last step. The first five steps work whether you do them yourself, hire someone, or use us.
Key takeaways
- Digitize before anything else. Analog tape degrades and the machines to play it are disappearing. The Library of Congress found its own 1970s preservation tapes suffering from sticky-shed syndrome, and reported a virtual cessation of manufacturing for analog tape media and analog tape machines (Library of Congress, via the National Archives).
- Capture a preservation master at 96 kHz and 24-bit in WAVE format, then work from a compressed copy (Library of Congress, via the National Archives).
- Expect automatic transcription to struggle. On 232 recorded interviews with older adults in Dutch long-term care, half of them in regional dialect, an off-the-shelf model produced a 48.3% word error rate, which fell to 24.3% only after fine-tuning on 34 hours of the same kind of audio (JAMIA).
- Transcribe verbatim, then edit. Oral history practice is to keep enough of the false starts and repeated phrases to show individual speech patterns, and to mark unclear passages rather than guess at them (Baylor University Institute for Oral History).
- Do not fill gaps with plausible invention. Mark them. A memoir that admits what the recording does not say is worth more than one that quietly makes it up.
- Under US practice, copyright in a recorded conversation sits with both the person speaking and the person recording, unless someone signed it over (Oral History Association). The rules differ by country.
- Keep the audio next to the book. The recording is a second artifact, not a raw ingredient you throw away once the text exists.
Take inventory before you touch anything
Get everything in one place and write a list before you play a single tape. You are looking for more than you think you have.
Analog tape. Standard cassettes, microcassettes from dictation machines and answering machines, reel-to-reel, and the audio buried on camcorder tapes (Video8, Hi8, MiniDV, VHS-C). Answering machine microcassettes are the ones families forget, and they are often where a voice survives when nothing else did.
Digital files you already have. Voice memos on a phone, WhatsApp voice notes, Zoom and FaceTime recordings from the pandemic years, Facebook Messenger audio, files copied off an old laptop.
Voicemail. Treat every saved voicemail as temporary. Where it actually lives depends on the phone and the carrier. Visual voicemail keeps a copy on the device, while older carrier voicemail keeps it on a server and deletes it on a schedule you do not control. Either way, the message is one lost phone or one carrier migration away from gone. Export each one to a real audio file today, before you plan anything else. On an iPhone, the share option in the voicemail list will send a message to Voice Memos, Files, or email. On Android it depends on the carrier's app, and if there is no export option, play it on speaker and record it with a second phone. A degraded copy that exists beats a clean one that got deleted.
Video with usable sound. A wedding tape where someone gives a twenty-minute toast about their childhood is an oral history with pictures attached.
For each item, write down the format, roughly when it was made, who is on it, and how long it is. That list is the spine of the whole project, and it is the thing that tells you whether you have a book or a chapter.
How do you get old audio off the tape?
The tape is the only part of this with an actual deadline. Everything else can wait a year. This cannot.
Magnetic tape does not fail gracefully. The Library of Congress documented that its own analog preservation copies made in the 1970s later suffered from sticky-shed syndrome, and that those deteriorating preservation copies had themselves to be reformatted. The same paper makes the obsolescence point plainly: on the audio side it reports a virtual cessation of manufacturing for analog tape media and analog tape recording devices (Library of Congress, via the National Archives). That was written in 2003. Sticky shed itself is a binder problem, where the layer holding the magnetic coating absorbs moisture, breaks down, and sheds residue onto the playback head. The 2015 ARSC guide, published with the Library of Congress, opens on exactly this: the audio legacy is at serious risk from media deterioration and technological obsolescence (ARSC Guide to Audio Preservation).
What this means in practice:
Play a fragile tape as few times as possible. Every pass costs you something. Capture on the first play, not the third.
If a tape squeals, sticks, or leaves powder on the heads, stop. That is the binder failing. Transfer studios treat it with low-temperature baking, which opens a window for a single transfer rather than repairing anything. This is the point to pay a transfer service rather than experiment on the only copy of your grandmother's voice.
Capture a preservation master and an access copy. The preservation master is uncompressed WAVE at 96 kHz and 24-bit, which is the specification the Library of Congress moved to for reformatting (Library of Congress, via the National Archives). It is bigger than you need and that is the point: you will never do this again. Then make an MP3 or M4A of each file to actually work from.
Do not clean up the master. Noise reduction, hum removal, and loudness normalization all belong on a copy. Filters that make speech clearer to a transcription tool also shave off the breath and room tone that make a voice sound like a person.
Name files so a stranger could sort them. Date first, then who and what: 1994-06-xx_grandma-ruth_kitchen-table_side-a.wav. Unknown dates get the decade and a note.
Back up in three places the same day. Local drive, second drive, cloud. The most common way these files are lost is a laptop, not a decision.
Doing the same job on the photographs and film in the same box is a parallel project with its own settings and pitfalls, and we covered it in how to digitize old family photos and videos.
How accurate is AI transcription on old family audio?
Better than most people expect on clean audio, and much worse than the marketing implies on the audio you actually have.
The honest benchmark comes from research rather than vendor pages. In a study of 232 recorded interviews in Dutch long-term care, half of them conducted in regional dialect, an off-the-shelf speech recognition model hit a 48.3% word error rate. Fine-tuning it on 34 hours of the same kind of audio brought the average down to 24.3%, with residents still the hardest group to transcribe. The authors attribute the difficulty to background noise, age-related speech changes, and regional dialect, and note that human transcribers could not always make out the audio either (JAMIA). The setting is not yours, but the conditions are: an unplanned recording, an older speaker, room noise, and an accent no model was tuned on.
Your box of tapes has every one of those problems and several more: tape hiss, a microphone across the room, two people talking over each other, forty years of generational slang, and place names no model has seen. Assume the first machine pass is a rough draft, not a transcript.
So the workflow is:
Run the machine pass first anyway. Even a transcript with errors at the rates above gives you searchable text and timecodes, which is what you need to find the good material across twelve hours of audio.
Ask for verbatim output, not "clean" or "smart" formatting. Clean transcription strips exactly the speech patterns you are trying to preserve. You can always cut a repeated phrase later. You cannot put back one you never captured.
Audit against the audio, in one pass, with headphones. Fix proper nouns first: names, towns, ships, regiments, employers, streets. These are what a model gets wrong and what a family notices instantly.
Use a convention for what you cannot make out. Baylor's transcribers type the closest approximation, underline the questionable portion, and add two question marks in parentheses. Where no guess is possible at all, they leave a blank line of roughly the right length and mark it the same way: "We'd take our cotton to Mr. _________(??)'s gin in Cameron" (Baylor University Institute for Oral History). Never let a guess enter the text unmarked. Six months later you will not remember which words were heard and which were assumed.
Keep bracketed sounds factual. Baylor's rule is to write (laughs) rather than (chuckles) or (guffaws), and (both talking at once) rather than (interrupts). Descriptions are interpretation, and interpretation belongs to the reader.
Turning fragments of old voice recordings into memoir chapters
This is the part almost no guide covers, and it is where turning voice recordings into a memoir stops resembling a normal writing project. The problem only exists because the recordings were not made as an interview. You do not have a life story. You have thirty pieces of one, out of order, with no index.
Index by timecode before you write a word. For each file, produce a line every couple of minutes: timestamp, topic, who is speaking, and a note if the passage is unusually good. A twelve-hour collection becomes a five-page index. That index, not the audio, is what you work from for the rest of the project.
Sort the index into a life spine. Childhood, family of origin, school, work, marriage, children, moves, war or migration, later life, beliefs. Now you can see the shape of what you have: three hours on the farm and the war, nine minutes total on forty years of marriage.
Let the chapters follow the audio, not the biography. The strong instinct is to write a complete life and pad the thin parts. Resist it. A book of eight chapters where the recordings are rich is better than eighteen chapters where ten are invented connective tissue. Uneven coverage is honest, and it is also more interesting to read.
Build chapters around one recording each where you can. A single sitting usually has its own arc. Chapters assembled from six different tapes tend to read like a summary rather than a person talking.
Put the fragments that fit nowhere in one section at the back. Call it what it is: a set of things they said. Half a page of one-paragraph pieces about a dog, a boss, a song. Nothing is lost by keeping them, and a fragment with no home is still in their voice.
Use a light narrator frame, in a different voice from theirs. One or two lines of italic before a chapter to place it in time ("Recorded in the kitchen on Ruth's eightieth birthday, with the grandchildren in the next room"). That frame is where the book can say what the audio does not, without ventriloquizing anyone.
If your box also holds letters and diaries, they can carry the years the recordings skip. We wrote a separate method for that material in what to do with old letters and journals from a deceased relative.
What do you do about the gaps you can no longer ask about?
There are exactly three honest moves, and one common dishonest one.
Corroborate. Documents settle dates, spellings, and sequence better than memory does: certificates, service records, ship manifests, employment records, parish and civil registers. If a sibling or cousin was on the same tape or in the same room, ask them. Two accounts that disagree are more useful than one that is unchallenged.
Mark the gap in the text. "She never said on any recording why the family left Bergen in 1954." That sentence is a real finding. It tells a grandchild what to go looking for, and it protects the rest of the book's credibility.
Ask the living generation now. The recordings you already have are fixed. The people who are still here are not. If this project has taught you anything, it is that the next interview should happen this month. Our list of unexpected ways to capture a parent or grandparent is the low-friction version, and the 90-minute interview structure is the thorough one.
The dishonest move is asking a language model to write the missing chapter in your grandmother's voice. It will do it, and it will be fluent, and the result is fiction wearing her name. Software is genuinely good at reorganizing what someone actually said. It is not a source for what they did not say. Keeping that line clean is the difference between a family record and a forgery.
Editing without erasing the voice
The edit is where written voice usually dies, and the fix is procedural rather than a matter of taste.
Cut fillers with a quota, not a rule. Baylor's transcribers type no more than two crutch words per occurrence per page, on the reasoning that a fully verbatim page is exhausting to read while a page with none has been sanded down into somebody else's prose (Baylor University Institute for Oral History).
Keep the false starts that carry personality. The same guide takes a deliberate middle course, leaving in enough repeated words and phrases to indicate individual speech patterns, and always keeping repetition that the speaker used for emphasis.
Keep their word order. People from a first language other than the one they are recorded in have a rhythm to their sentences. Correcting it into idiomatic prose is the single fastest way to make a book stop sounding like them.
Read every chapter aloud against the tape. Play two minutes, read the corresponding page, play it again. If the page sounds like a different person, the page is wrong. This test beats every style rule, and it is the reason we tell people to draft from the transcript with the audio queued up, which is the same discipline described in how to preserve a parent's voice in a book.
Fix only what obstructs. Punctuation, paragraphing, and order of sections are fair game. Vocabulary is not.
Who is allowed to publish a recording you did not make?
Worth settling before you print, especially if the book leaves the family.
Copyright in an interview usually sits with both parties, the person speaking and the person who recorded them, unless one of them signed it over. Oral history archives handle this with a legal release from both, and they treat access restrictions as the narrator's to set (Oral History Association). That is the US position, and the details differ by country, so check locally before anything leaves the family. For a private family book that stays in the family, this is mostly theoretical. For anything published, sold, or posted, it is not.
Two practical steps. Get written permission from anyone living who is audible on the recordings and quoted at length, including the person who did the recording. And ask the room what should stay in the room: a diagnosis, an affair, an estrangement, someone's paternity. The archival answer is to produce two editions, a family one and a sealed one, and it works just as well at a kitchen table as in a repository.
Printing it, and keeping the audio next to the book
The book is not the only artifact. The recording is the other one, and it is the one that cannot be recreated.
Print at least three copies and put one with a relative in a different house. Include a short note on provenance at the front: which recordings the book was built from, when they were made, who transcribed them, and where the audio files now live. That page is what makes the book usable to someone in 2075.
Put the audio somewhere a family member can actually reach it, not just a drive in your desk. A cloud folder shared with everyone, a QR code printed in the back of the book, or both. And keep the preservation masters even though nobody will ever open them, because every future format conversion should come from those and not from the MP3.
If you want the numbers on what production actually costs, we broke them down in how much it costs to make a memoir book, and the tools for each step in this workflow are in the 17 tools we recommend.
What this looks like with Yourtale
We should be precise about where we fit, because it is not the whole of this article.
We interview living people. We run the conversation over voice across a few sessions, keep the audio, transcribe it verbatim, and draft chapters directly from the transcript. The chapters are drafted by AI from what was actually said, not written by a ghostwriter, and the only human editor is you: no Yourtale employee reads or shapes your book. The family reviews and approves every chapter before it prints. The hardcover is $199 at the founding rate.
Bringing recordings you already have into a Yourtale book is something we currently do by hand rather than through a button in the product. If that is your situation, talk to us before assuming it is automatic.
And if the person in your recordings is gone, the honest advice is the whole method above, which does not require us. Digitize this month. Transcribe verbatim. Index by timecode. Build chapters around what is there and mark what is not. Then go record whoever is still here, because that is the only part of this you can still change. Join the waitlist if you want us to run that interview.
Frequently asked questions
Can I turn old cassette recordings of my parent into a book?
Yes, and the sequence matters. Digitize the tapes first, because analog tape degrades and playback machines are getting scarce. Transcribe the digital files verbatim with an automatic tool, then audit that transcript against the audio to fix names and unclear passages. Index every recording by topic and timecode, group the material into chapters that follow the coverage you actually have, edit lightly against the audio, and print. The limiting factor is not technology. It is that you cannot ask follow-up questions, so the book has to be built around fragments.
What is the best way to transcribe old family audio recordings?
Run an automatic pass first for searchable text and timecodes, then audit it against the audio with headphones. Ask for verbatim output rather than clean or smart formatting, since cleaning strips the speech patterns you want to keep. Expect substantial errors on degraded or accented audio: a study of 232 recorded interviews with older adults in Dutch long-term care, half of them in regional dialect, found a 48.3% word error rate from an off-the-shelf model, falling to 24.3% after fine-tuning on similar material. Fix proper nouns first, and mark anything you could not make out rather than guessing.
How do I digitize old cassette tapes and voicemails?
For cassettes, use a working deck or a USB cassette player, capture an uncompressed WAVE master at 96 kHz and 24-bit, and make a compressed copy to work from. Play fragile tapes as few times as possible, and send squealing or sticky tapes to a transfer service instead of risking the only copy. For voicemails, export each message to an audio file immediately using the share option in your phone's voicemail list. Where the message lives depends on the phone and the carrier, and carrier-side voicemail is deleted on a schedule you do not control.
Should I let AI write the parts my recordings do not cover?
No. Software is good at reorganizing what someone actually said, and using it to invent what they did not say produces fluent fiction under a real person's name. The three honest options are to corroborate the gap with documents or other relatives, to state plainly in the text that the recordings never covered it, or to interview someone still living who knows. A memoir that marks its gaps is more credible than one that hides them.
How many hours of recordings do I need for a memoir?
There is no minimum, and the useful measure is coverage rather than duration. Five hours spread across childhood, work, and family will support a fuller book than twenty hours about one subject. Index what you have by topic before deciding: most families discover their coverage is very uneven, with hours on a few subjects and minutes on decades. Let the chapters follow that unevenness instead of padding the thin parts.
Who owns the copyright to an old recording of a family member?
Under US practice, copyright in a recorded interview belongs to both the person speaking and the person who did the recording, unless one of them transferred it in writing. Archives handle this with a signed release from both parties and treat access restrictions as the narrator's decision. The rules differ by country, so check locally. For a private book that stays inside the family this is rarely tested, but for anything published, sold, or posted online you should get written permission from everyone living who is audible and quoted.
Sources cited above
- Carl Fleischhauer, Library of Congress, "Reformatting Audio and Video and the Motivations for a New Approach", Preservation Conference, National Archives at College Park, March 2003: sticky-shed syndrome in the Library's own analog preservation tapes, format obsolescence, and the 96 kHz / 24-bit WAVE reformatting specification. Retrieved 2026-09-08.
- ARSC Guide to Audio Preservation, Council on Library and Information Resources with the Association for Recorded Sound Collections and the Library of Congress, 2015: recorded-sound heritage at risk from media deterioration and technological obsolescence. Retrieved 2026-09-08.
- Journal of the American Medical Informatics Association, "The development of an automatic speech recognition model using interview data from long-term care for older adults", 2023: 48.3% word error rate before fine-tuning, 24.3% after, across 232 interviews and 34 hours of audio. Retrieved 2026-09-08.
- Baylor University Institute for Oral History, Style Guide: A Quick Reference for Editing Oral History Transcripts, revised May 2018: crutch-word limit, false-start policy, unintelligible-passage notation, and non-editorializing bracketed sounds. Retrieved 2026-09-08.
- Oral History Association, Archiving Oral History: Manual of Best Practices, adopted October 2019: legal release, shared copyright between interviewer and narrator, preservation of the audiovisual original, and narrator-set access restrictions. Retrieved 2026-09-08.