Methodology · how a reading works
How Resonate measures a voice.
The bands on this page are read from the same tables the scoring pipeline uses, so they can’t drift from the product. Measured voices last updated 2026-07-26.
A Resonate reading turns about 20 seconds of speech into 19benchmarked acoustic measurements and six perception scores. The pipeline tracks pitch roughly 100 times a second, measures pacing, melody, loudness, voice quality and resonance directly from the signal — not from a transcript, and not from a language model’s impression — and scores each measurement against a research-derived band: an optimal range and a wider acceptable range. The six headline scores — Conversational, Public speaking, Authority, Warmth, Likeability and Vocal attractiveness — are weighted reads of those same acoustics, scored against listener-perception research on a 0–10 scale.
Two things hold on every surface. Scores are never inflated or blurred — an honest 5.4 with a clear path up is worth more than a flattering 8.2 you can’t trust. And the free reading stores zero recordings: the audio is deleted the moment it’s analysed.
What do the six dimensions measure?
Resonate groups its 19 benchmarked measurements into six dimensions — the trainable shape of a voice. Each dimension is fed by specific acoustic measurements, so a weak reading always points at something concrete to work on.
Melody
How much your voice moves
Measured from
Pitch variability · Pitch range · Pitch dynamism (PDQ)
Cut-through
Whether your words actually reach people
Measured from
Presence band (2-4 kHz) · Spectral tilt (alpha ratio) · Voice clarity (HNR)
Pace & silence
Speed, rhythm, and the confidence to pause
Measured from
Speaking rate · Articulation rate · Syllable rate · Pause ratio · Average pause length · Phrase length
Control
Breath support and steadiness
Measured from
Pitch steadiness (jitter) · Volume steadiness (shimmer) · Vocal fry (creak)
Warmth
The richness people relax into
Measured from
Chest band (0-500 Hz) · Average pitch
Words
Fillers, clarity of diction
Measured from
Filler words · Intelligibility (ASR)
Two loudness measurements — Volume dynamics and Loudness variability — are benchmarked and scored the same way but sit outside the six training dimensions, which is why the full table below has 21 bands.
What are the benchmark bands?
Every reading is scored against the 21 benchmark bands below: for each acoustic measurement, a research-derived optimal range and a wider acceptable range — the same numbers for everyone, with no adjustment for reputation. Measured against these same bands, Barack Obama’s pause ratio (21.2%) reads inside the 18–35% optimal band, and Winston Churchill’s 15.5 st pitch range reads above the 7–12 st optimal band — the measured-voice library shows every readout in full.
Bands split by sex only where physiology genuinely splits them: average pitch carries a separate female band, shown below. Everything else applies to every voice.
| Measurement | Unit | Optimal | Acceptable | What it measures |
|---|---|---|---|---|
| Pitch & melody | ||||
| Average pitch | Hz | 100–150female: 180–230 | 80–185female: 160–260 | Mean fundamental frequency. Typical adult male ≈ 85-155 Hz, female ≈ 165-255 Hz. |
| Pitch variability | st | 3–6.5 | 1.8–9 | Standard deviation of pitch in semitones. Under ~1.5 st reads as monotone; engaging/charismatic speakers typically sit above 2.5 st. |
| Pitch range | st | 7–12 | 4–18 | Span between 10th and 90th percentile pitch, in semitones (12 st = one octave). Typical reading ≈ 7-10 st; engaging speakers ≈ 8-12 st. |
| Pitch dynamism (PDQ) | — | 0.15–0.38 | 0.08–0.55 | Pitch SD ÷ mean pitch (Hincks). Sex-normalised liveliness: ~0.11 sounds flat, 0.2+ sounds lively. Band is calibrated to agree with the pitch-variability (st) band. |
| Pacing & pauses | ||||
| Speaking rate | wpm | 140–160 | 115–190 | Words per minute including pauses. Conversational sweet spot for engaging delivery. |
| Articulation rate | wpm | 160–220 | 120–270 | Speed while actually talking (pauses excluded). High articulation + good pauses = dynamic but clear. |
| Pause ratio | % | 18–35 | 8–50 | Share of total time spent silent. Great speakers pause deliberately — around a fifth to a third of the time. |
| Average pause length | s | 0.4–1 | 0.25–1.8 | Typical pause duration. 0.5-1.0 s pauses aid comprehension; much longer reads as losing your place. |
| Syllable rate | syl/s | 4–5.5 | 3–6.8 | Estimated syllables per second while speaking (pauses excluded). Fluent adult speech runs 4-5.5. |
| Phrase length | words | 5–12 | 2–22 | Average words between pauses. Charismatic speakers chunk ideas into short phrases (Jobs vs Zuckerberg studies). |
| Dynamics | ||||
| Volume dynamics | dB | 8–18 | 4–28 | Spread between quiet and loud moments (5th-95th percentile of frame intensity). Flat volume = flat interest. |
| Loudness variability | dB | 3.5–8 | 2–12 | Moment-to-moment intensity variation. Correlates with perceived enthusiasm. |
| Voice quality | ||||
| Voice clarity (HNR) | dB | 12–30 | 6–40 | Harmonics-to-noise ratio over connected speech. Higher = clearer, more resonant tone; low values sound breathy or under-supported. |
| Pitch steadiness (jitter) | % | 0–2.5 | 0–6 | Cycle-to-cycle pitch instability measured over connected speech (runs higher than the 1.04% sustained-vowel clinical threshold; phone processing adds more). |
| Volume steadiness (shimmer) | % | 0–9 | 0–20 | Cycle-to-cycle amplitude instability over connected speech (the 3.81% clinical threshold applies to sustained vowels; running speech and phone AGC read far higher). |
| Vocal fry (creak) | % | 0–8 | 0–25 | Share of voiced frames in the creaky low register — usually sentence endings running out of breath. A little is natural; a lot reads as low-energy. |
| Spectral balance | ||||
| Presence band (2-4 kHz) | % | 6–20 | 2–35 | Energy share in the 2-4 kHz 'speaker's formant' region — this is what cuts through rooms and microphones. |
| Chest band (0-500 Hz) | % | 35–75 | 15–90 | Low-frequency energy share — warmth and chest resonance. Too little sounds thin/nasal, too much sounds boomy or muffled. |
| Spectral tilt (alpha ratio) | dB | -22 to -8 | -32 to 0 | Balance of 1-5 kHz vs 50-1000 Hz energy. Higher (less negative) = brighter, more projected; lower = darker, softer. |
| Language | ||||
| Filler words | /min | 0–2 | 0–6 | 'Um', 'uh', 'like', 'you know'… per minute. Under ~2/min goes unnoticed; above ~4/min distracts. |
| Intelligibility (ASR) | — | 0.8–1 | 0.5–1 | How confidently a strong speech recogniser decodes your words (Whisper avg token probability). A direct proxy for 'easy to understand'. |
Units: wpm = words per minute; st = semitones (12 st = one octave); HNR = harmonics-to-noise ratio. The two language measurements need a transcript, so they’re read only where transcription runs; everything else comes straight from the signal.
Does the same recording get the same scores?
Yes — exactly the same. In July 2026 we uploaded one 61-second clip five times through the deployed production pipeline, the same path every free reading takes: all five responses came back byte-identical once the response timestamp was stripped — that server-clock field is the only thing that varied — and two further runs at full depth returned every numeric value — every metric, every point of the pitch contour — exactly equal. The analysis is deterministic, with no run-to-run noise, so any difference you see between takes is a difference in your voice, not in our measurement.
Re-compressing a recording changes the audio itself, slightly. Re-encoding that same clip from 128 kbps to 64 kbps MP3 moved the raw acoustic measurements by at most about 5.5% on substantive metrics (most moved under 1%), each of the six headline scores by at most 0.3 points on the 0–10 scale, and the overall score by 0.1 on its 0–100 scale. A single fine-grained metric sitting near a scoring boundary can move more — the largest single-metric score shift in that test was about 13 points on that metric’s 0–100 scale.
That is exactly what was tested, and no more: one clip, one codec pair, on the pipeline as deployed in July 2026. It shows the deployed pipeline adds no noise of its own — not that future versions of the scoring will match today’s. If the bands or the scoring change, that arrives as an announced change, not silent drift.
What doesn’t a reading capture?
The words. Resonate measures how a voice sounds, from the acoustics alone — it doesn’t judge the argument, the content, or the room you said it in, and a strong reading doesn’t make a weak point true. It is not a medical, clinical or diagnostic tool, and it doesn’t identify who is speaking.
The recording matters. Podcast and interview compression can flatten dynamics readings; narrow phone-line bandwidth can dull clarity; background noise can bury real pauses, so pacing reads busier than it was spoken. Where a source recording biases a reading we say so on the page — and when a measurement can’t be made reliably, it’s flagged as unread rather than rendered as a fake zero.
The measured-voice library is run without transcription: acoustic measurements only, so transcript-dependent readings such as filler words are absent there, and each voice’s overall score weighs the acoustic set. Comparisons are based on Resonate's acoustic analysis of publicly available speeches. No endorsement or affiliation is implied. Every voice page documents its source recording and its condition.
Where a reading points somewhere specific, the guides go deeper: speaking pace, pause and silence, the authoritative voice and the attractive voice. The rest of the common questions live in the FAQ.