Audiobook length calculator: estimate runtime from your manuscript

Word count gives a first-pass audiobook runtime. An editable speaking rate and a timed sample make the estimate more useful; the mastered files supply the final duration.

EarnDraft Team · September 26, 2026 · 21 min

Headphones, an open book, a paper waveform, and a stopwatch.
Editorial illustration created for EarnDraft with AI image generation.

Interactive planning tool

Estimate your audiobook runtime

Enter your manuscript length and an assumed narration pace. This is planning arithmetic, not measured EarnDraft audio data or a finished recording.

Estimated finished time

5 hr 23 min

At an illustrative 130–190 words per minute, the same manuscript spans about 4 hr 23 min to 6 hr 25 min. Your actual narrator, pauses, dialogue, and production choices will change the result.

How editable speaking speed changes the estimate
  1. Slower scenario140 words per minute
  2. ACX human planning benchmark155 words per minute
  3. Faster scenario170 words per minute
In this guide

An audiobook’s finished length affects the listener’s expectations, production schedule, file plan, and store metadata. Word count gives a useful first estimate, but it is not a recording. Narrators pause, dialogue changes pace, technical terms slow delivery, and opening credits add time. A calculator should therefore show its assumptions and a range. It should invite an author to measure a sample, then replace the estimate with the actual mastered duration before publication.

This page gives the arithmetic, worked examples, a sample-timing method, chapter planning, and common reasons the estimate changes. The central benchmark is the one ACX currently gives for human performers: about 9,300 words per finished hour, or 155 words a minute. It is an approximation for planning, not a promise for every voice, language, genre, or AI narrator. EarnDraft helps organize and draft book text at /create; use an audio production and playback workflow to measure the actual narration.

Inkfluence AI’s own audiobook length calculator reports a mean of 164.4 words per minute across 8,015 finished chapters, measured on September 15, 2026. Its published FAQ also reports 1,552 hours of audio and says the middle 80% of chapters ran between 138 and 192 words per minute. These are Inkfluence’s first-party measurements of its own narrated output, not an independent cross-platform benchmark and not EarnDraft usage data. The range is useful evidence that a single fixed rate can miss substantial variation. For a new project, use an editable assumption and then time the actual chosen voice.

Calculate a first-pass runtime

The basic formula is spoken words ÷ spoken words per minute = narration minutes. Divide by 60 for hours. Count the words that will actually be spoken, not the whole manuscript file. Exclude page numbers, footnotes not intended for audio, tables replaced by a spoken summary, and production notes. Include opening and closing credits, spoken chapter titles, introductions, and any other material the listener hears. A novel’s text may be close to the spoken script; a workbook with many tables may need a separate audio adaptation.

Spoken words At 140 words/min At 155 words/min At 170 words/min
10,000 1 h 11 m 1 h 5 m 59 m
25,000 2 h 59 m 2 h 41 m 2 h 27 m
50,000 5 h 57 m 5 h 23 m 4 h 54 m
75,000 8 h 56 m 8 h 4 m 7 h 21 m
100,000 11 h 54 m 10 h 45 m 9 h 48 m

The table is arithmetic, not a measured distribution. The slower and faster rates are scenario assumptions used to show sensitivity. ACX’s 155 words a minute is an average planning benchmark for performers; it does not establish a specific AI narrator’s speed. A highly dramatized reading may take longer. A brisk instructional narration may be faster, but clarity should decide pace. For a first planning range, calculate at several rates, then time a representative excerpt in the voice and style you intend to release.

If you need a spreadsheet, use =spoken_words / words_per_minute / 60 for hours. To display hours and minutes, calculate total minutes first, then take whole hours and the remainder. Do not round each chapter before adding them; round the total at the end. For example, 64,000 spoken words at 155 words a minute is about 412.9 minutes, or 6 hours 53 minutes before extra pauses and credits. A change to 145 words a minute makes it about 7 hours 21 minutes. The speed assumption creates roughly half an hour of difference in this example.

Measure your own narrator with a representative sample

The strongest estimate uses audio from the actual production workflow. Choose an excerpt that resembles the book: dialogue and narrative for fiction, definitions and examples for a technical guide, or reflective prose for a memoir. Count only spoken words in the excerpt. Record or synthesize it using the planned narrator, pronunciation settings, and pacing. Listen for intelligibility and emotional fit before treating its duration as a useful rate. A too-fast sample is not a desirable benchmark simply because it yields a shorter book.

Sample field Example entry Why record it
Passage ID Chapter 4, pages 8–12 Reproducibility
Spoken words 1,240 Formula input
Final sample duration 8 m 35 s Includes natural pauses
Observed rate About 144 words/min Personal estimate
Content type Dialogue with three names Explains pace
Corrections Two name pronunciations Production work still needed

For the fictional sample above, 8 minutes 35 seconds is about 8.58 minutes. Dividing 1,240 by 8.58 gives about 144 spoken words a minute. A 70,000-word script at that rate would be approximately 8 hours 6 minutes before credits and chapter transitions. This is a scenario, not evidence that any particular narrator runs at 144. Time at least one other kind of passage if the book varies. A single quiet scene may underestimate the speed of action chapters or overestimate a chapter dense with invented names.

Use the finished sample duration. Raw recording time includes retakes and silence that will be edited out; production labor is a separate number. Conversely, a hastily synthesized file may lack pauses you will add for the final listener experience. Once you settle the narration style, measure the edited sample and update the estimate. If the author changes the script during audio proofing, update the spoken word count too. Version the count beside the sample so old numbers do not circulate as current metadata.

Build a chapter-level estimate

Total runtime does not reveal whether one chapter will create an unwieldy file or a long listener session. Count each chapter separately and add transition time. ACX’s current audio submission requirements say each submitted file should be no longer than 120 minutes and should contain one chapter or section. A chapter approaching that ceiling at your measured rate needs a production plan. The technical requirement does not by itself mean AI audio is eligible for ACX; its current rules require human narration unless otherwise authorized.

Section Spoken words At 155 words/min Production note
Opening credits 90 Under 1 min Separate file where required
Prologue 2,600 17 min Test tone and names
Chapter 1 6,800 44 min Normal chapter
Chapter 2 9,200 59 min Several dialogue voices
Chapter 3 14,500 94 min Long but below 120-minute example limit
Closing credits 110 Under 1 min Separate file where required

These are invented numbers for planning. Apply the current rules of your actual destination store and production tool. If a section is very long, splitting it may improve navigation even when it fits a technical limit. Do not split at a random midpoint: choose a scene or argument boundary and name the new section clearly. If chapter titles are spoken, include them in the script and estimate. If music or long room tone is included, a word-count formula cannot predict it; add measured durations separately.

For a book with many short chapters, credits and transition pauses can form a larger share of the total. A thirty-day devotional may contain the same total words as a novella but more section openings. A table-heavy guide may require spoken explanations that are longer than the printed labels. A children’s book may be read more slowly, with deliberate pauses for illustrations or page turns. Use the formula to plan; use a sample to calibrate.

Distinguish runtime from production time

A six-hour audiobook does not necessarily take six hours to make. Human narration involves preparation, recording, retakes, editing, proof listening, mastering, metadata, and upload. AI narration can generate audio quickly but still needs script preparation, name pronunciation, listening checks, corrections, file assembly, and store review. The number of finished hours is one input to a schedule, not the schedule itself.

Work phase Output Why it can take longer than playback
Script adaptation Spoken edition Tables, notes, and links need audio treatment
Pronunciation prep Name and term guide Consistency across chapters
Recording/generation Raw chapter audio Retakes or regeneration
Proof listening Error log Human must hear the actual output
Edit and master Final files Repairs, levels, silence, encoding
Package Cover, credits, metadata Store-specific requirements
Review Approved listing External processing and corrections

Plan time for at least one complete listen of the released edition. Spot-checking only the first chapter can miss a wrong recurring name or a missing paragraph late in the book. A proof listener should compare audio against the final script, mark timecodes, and classify defects as pronunciation, omission, duplication, pacing, voice inconsistency, or technical noise. Listen on ordinary headphones and a phone speaker; a file that sounds fine on studio speakers may be tiring or unclear in daily listening conditions.

Do not confuse file encoding with listener quality. A file can meet sample-rate and loudness specifications yet contain wrong words or an unnatural performance. Technical validation is a gate, not a review of storytelling. If a store accepts uploaded AI narration, check its current file and disclosure rules separately from the audio content review. The guide where to sell an AI audiobook covers distribution decisions.

Adjust the estimate for the kind of book

The manuscript’s form changes spoken length. Dialogue may require slight pauses between speakers, and an actor may perform action and reflection at different speeds. Dense nonfiction may need slower articulation, especially around acronyms, equations, or lists. A cookbook has ingredients and measurements that are awkward to read as continuous prose. A poetry collection uses deliberate silence. These are reasons to sample the actual material, not universal speed multipliers.

Book form What word count misses Better sample
Novel Dialogue beats and emotional pauses Scene with dialogue and narration
Epic fantasy Invented names and pronunciation Passage with recurring names
Business guide Bullets, acronyms, diagrams Dense explanatory section
Memoir Voice and reflective pacing Emotional narrative section
Devotional Many short openings and pauses Several complete entries
Children’s book Illustration and page-turn rhythm Full story at intended delivery
Workbook Tasks not naturally spoken Audio-adapted exercise

For fiction, keep a name pronunciation sheet. It reduces repeated corrections and can make measured sample pace representative of later chapters. A narrator who stumbles on every invented place name will not maintain the planning speed. In a series, reuse the pronunciation sheet and document the chosen voice and pace. See writing a book series with AI for continuity planning and fiction audiobook production for performance review.

For nonfiction, decide whether the audio edition is a literal reading or an adaptation. A chart titled “Comparison of twelve routes” may be clear on a page and incomprehensible when spoken cell by cell. Rewrite the point of the chart as a brief narrative and provide a supplemental PDF if the store supports it. Count the adapted audio script, not the print layout. Put URLs in show notes or supplemental material where appropriate rather than reading an unwieldy link aloud.

How speed changes cost and file planning

Finished hours often influence narration quotes, editing effort, and listener price categories, but each service sets its own terms. Avoid multiplying an estimated length by a copied “industry rate” and treating it as a guaranteed invoice. Obtain current quotes or vendor prices for the exact production scope. Ask whether the quote includes proof listening, corrections, mastering, credits, and store-ready files. A low per-hour recording price may omit the work that makes the audio usable.

At a constant encoding rate, a longer audiobook tends to create larger files. However, file size also depends on codec, bitrate, channels, and how chapters are packaged. Do not calculate upload size from word count alone. Render a representative chapter in the required format, note its actual size, and scale roughly from duration while preserving a margin for credits and metadata. Check the destination store’s current size and encoding requirements. A single merged MP3 may be convenient for direct download but chaptered files can improve navigation and meet store upload rules.

The author should plan a quality budget as well as a production budget. A longer book means more opportunities for pronunciation errors, repeated sentences, chapter-order mistakes, and voice drift. A cheap, fast generation method can create more review work than a carefully prepared script. A measured pilot chapter helps estimate correction effort. Log defects per finished hour for your own project; do not present that private ratio as a market statistic. If the first chapter needs extensive repair, fix the script or pronunciation workflow before processing the rest.

Worked scenarios: same words, different outcomes

Imagine two 80,000-word books. The first is a contemporary novel with straightforward names and long narrative paragraphs. The second is a fantasy novel with many invented names, short dialogue exchanges, and chapter epigraphs. At ACX’s approximate human planning rate of 155 words a minute, both start near 8 hours 36 minutes of spoken narration. That is only a common starting point. The fantasy book may need longer pauses and pronunciation correction. The contemporary novel might be performed more slowly for emotional texture. Neither author can claim the actual runtime from word count alone.

Now imagine a 35,000-word instructional book. The print edition contains tables, screenshots, and checklists. If 4,000 words of captions and raw table cells will not be spoken, the audio script begins around 31,000 words. If the author replaces those visuals with 2,000 words of clear explanation, the final spoken script is 33,000 words. At 155 words a minute, that is about 3 hours 33 minutes before credits and pauses. The original print word count would have suggested about 3 hours 46 minutes. The difference is not large in this example, but the adapted version may be much more understandable.

Scenario Print words Spoken script Planning rate Arithmetic result
Contemporary novel 80,000 80,000 155 wpm About 8 h 36 m
Fantasy novel 80,000 80,000 155 wpm Same arithmetic, more performance uncertainty
Instructional print book 35,000 33,000 155 wpm About 3 h 33 m
Short audio-first guide 18,000 18,500 with credits 145 wpm About 2 h 8 m

All figures in the table are invented inputs and arithmetic outputs. They illustrate the method, not observed runtimes or an expected genre speed. A calculator that displays a single precise minute without its rate can imply more certainty than the data supports. Show the assumed speed, spoken word count, and whether credits are included. When the actual mastered edition exists, use the measured duration in the store listing.

A transparent calculator specification

If you build your own calculator, let the author enter manuscript words, excluded words, audio-only additions, and rate. Show the resulting spoken count and runtime range. Include a reset or example button, but do not silently change the rate based on genre without evidence. If you offer presets, label them as scenarios. The most valuable field may be “measured sample words and duration,” which computes a custom rate and makes the estimate specific to the project.

Input Validation Output effect
Manuscript words Nonnegative whole number Starting total
Excluded print-only words No more than manuscript count Subtract
Added audio words Nonnegative Add
Scenario rate Positive words/minute Runtime estimate
Sample words Positive when sample supplied Custom rate numerator
Sample seconds Positive when sample supplied Custom rate denominator

The formulas are simple: spoken = manuscript − excluded + added; sample rate = sample words ÷ (sample seconds / 60); runtime minutes = spoken ÷ chosen rate. Display a sensible rounded result such as “about 6 hours 50 minutes” and a note that pauses, performance, and final edit alter it. If the sample rate is implausibly high or low, prompt the user to check whether they counted only spoken words and used the final audio segment. Do not silently clamp the number and hide an input error.

An accessibility consideration: the result should be visible as text, not only as a colored gauge. A small bar chart can compare slow, reference, and fast scenarios, but the numbers and assumptions must appear in a table too. A screen reader should encounter labels and units. On mobile, a simple form and clear result are more usable than a dense dashboard. This guide’s table can serve as a no-script backup if an interactive calculator is unavailable.

Prepare audio metadata after measuring the release file

Stores may require duration metadata. Google Play Books’ content upload help explains how audiobook files are uploaded and ordered; its automated feed documentation describes duration metadata. Use the final mastered files to compute the value. Sum actual file durations, including credits, and verify the player’s displayed length after upload. A mismatch between store metadata and audio can confuse buyers and may indicate a missing chapter or duplicate file.

Keep a release manifest with section name, filename, duration, checksum if helpful, script version, and approval status. A manifest makes it easier to confirm that the audio package has the intended order. It also supports corrections: if one chapter changes, you can identify its file and the new total duration. For direct sale, include a clear read-me that says whether the buyer gets one file, chaptered files, or both. Do not promise a specific runtime in marketing while still making major edits.

Compare two candidate voices without rushing the story

When evaluating narrators, use the same passage and script version for each. Measure finished duration, but listen for comprehension, character distinction, and fatigue. A voice that reads 15% faster may shorten the file and production cost; it may also flatten dialogue or make technical material difficult to follow. The author should pick the voice that fits the book, then use its pace to update the runtime forecast. Do not select a voice solely because its shorter file fits a desired store category.

Review dimension Voice A Voice B Decision question
Sample duration Record actual result Record actual result Is pace comfortable?
Name pronunciation Log errors Log errors Can errors be fixed consistently?
Dialogue clarity Listen without script Listen without script Are speakers and beats understandable?
Emotional fit Note reader response Note reader response Does tone support the genre?
Long-session comfort Listen for 20–30 minutes Listen for 20–30 minutes Would this voice sustain the book?

Use at least one passage with the book’s hardest terms and one with its typical prose. A polished opening scene alone can hide weaknesses later. If the production tool allows pronunciation dictionaries or pace controls, set them before timing the sample. Changing those settings after estimating can shift every chapter. If a human narrator is hired, discuss pacing as a creative choice rather than demanding a specific words-per-minute number. A narrator can pause meaningfully where the story requires it.

Record the decision in a short production brief: narrator or voice ID, script version, sample passages, measured rate, pronunciation choices, and reasons for selection. For a series, this becomes a continuity artifact. If a future volume changes tool or performer, compare a shared passage and document the difference. The goal is not identical seconds per thousand words; it is a recognizable listening experience.

Plan uncertainty rather than hiding it

A runtime estimate can be expressed as a low, central, and high planning scenario. For a 60,000-word spoken script, the table above gives about 5 hours 53 minutes at 170 words a minute, 6 hours 27 minutes at 155, and 7 hours 9 minutes at 140. Those scenarios are arithmetic; they are not probability bounds. The author should not say there is a certain percent chance that the finished book falls in that range. Use the range to plan review time, file structure, and a realistic production schedule.

Uncertainty Direction of effect How to reduce it
Script still changing More or fewer words Freeze a spoken-script version
Voice not selected Different pace Measure candidate samples
Many technical terms Slower delivery, more retakes Pronunciation sheet and hard passage
Long pauses or music Adds duration Measure and add separately
Print-only content Spoken script may shrink or grow Adapt the text first
Corrections after mastering Total can shift Recompute from final files

It is helpful to distinguish precision from accuracy. The formula may calculate 6 hours, 27 minutes, and 5 seconds, but an unmeasured performance does not justify the seconds. Display rounded minutes and explain the rate. A decision such as “Will this chapter fit under a two-hour file limit?” may require a larger safety margin and a timed pilot. A decision such as “Should the cover say an eight-hour listen?” should wait until the final audio exists.

Use a chapter manifest to track variance. After finishing the first three chapters, compare actual durations with their predicted durations. If each is systematically longer, update the remaining forecast. If only dialogue-heavy chapters run long, classify chapters by material and sample accordingly. This turns the calculator into a living production instrument rather than a one-time search result. It also helps schedule proof listening: a revised nine-hour projection means nine hours of playback to review, plus time to investigate defects.

A handoff worksheet for a producer or collaborator

An author may know the manuscript count while a producer needs a clean, spoken script. Hand off a worksheet with the edition identifier, chapter list, exclusions, audio-only passages, narrator choices, and pronunciation guide. Note the target stores, because file organization and credits may vary. The worksheet should identify unresolved decisions instead of burying them in the manuscript. A producer can then estimate time from a sample and return a quote tied to clear scope.

Field Author provides Producer verifies
Script edition Date and version Matches supplied files
Spoken word count Per chapter and total Counts accepted after adaptation
Visual material Tables, charts, captions Spoken treatment or supplemental file
Names and terms Pronunciation notes Sample playback
Delivery Store and direct formats Current specs and eligibility
Credits Exact names and order Final spoken file
Quality process Expected listen and corrections Defect reporting path

This worksheet also prevents a common misunderstanding about a “finished hour” quote. Ask whether the price covers only performance or also editing, proofing, mastering, revisions, and file delivery. Runtime alone cannot answer that question. If using an AI voice, ask who listens to the complete output and who repairs mispronunciations. The labor does not vanish; it moves from microphone performance toward preparation and quality control.

Frequently asked questions

How many hours is a 50,000-word audiobook?

At ACX’s human narration planning benchmark of about 9,300 words per hour, 50,000 words is about 5 hours 23 minutes. At 140 words per minute it is about 5 hours 57 minutes; at 170 it is about 4 hours 54 minutes. The actual finished file determines the real length.

How many words are in one audiobook hour?

ACX says many human performers narrate about 9,300 words per finished hour on average. That is a planning benchmark, not a fixed rule. A project-specific sample gives a better rate for its voice, language, and material.

Is AI narration faster than human narration?

Generation time and listening speed are different. AI may produce files quickly, but the playback pace is set by voice and production choices. Time a sample from the actual tool and review it for clarity before estimating the full book.

Should I count chapter titles and credits?

Yes, if they will be spoken. Count the final audio script, including spoken front and back matter. Add measured music or silence separately because a word count cannot predict it.

Does a long chapter need to be split?

Check the rules of the destination store. ACX’s current audio submission requirements limit each file to 120 minutes and require one chapter or section per file. Even below a technical limit, a natural section break may improve navigation. Eligibility and format requirements are separate from runtime.

Can I use the estimate as the store’s duration?

Use it for planning. Once the audiobook is mastered, measure the real files and provide their actual total in metadata and the listing. Revise the number if a corrected edition changes duration.

Do fiction and nonfiction use different rates?

Their performance can differ, but a fixed genre multiplier is a poor substitute for a sample. Dialogue, technical terms, poetry, tables, and adaptation choices all affect pace. Time a representative passage from the actual script.

Sources and next step

The arithmetic uses the ACX narration planning benchmark and references its current audio submission requirements for chapter files. Google Play Books upload guidance illustrates why final file order and metadata matter. Check these primary sources again before production or submission because store rules change.

Count your spoken script, calculate a range, then record or generate a representative sample and replace the assumed speed with your observed one. Once the full audio is edited, listen through it and publish the measured duration. If you still need a book structure, start at /create; if you need a distribution plan, read where to sell an AI audiobook.

Written by the EarnDraft team. We make software for drafting, editing, exporting, and sharing short ebooks. Check changing platform rules at the linked primary source before publishing. About EarnDraft.

Sources and further reading

Put the idea into a practical book

Word count gives a first-pass audiobook runtime. An editable speaking rate and a timed sample make the estimate more useful; the mastered files supply the final duration.

Plan your book →