The Power of Voice: What Executives’ Voices on Earnings Calls Revealed, and Why the Method Is Disputed

By Noam Shemla · The Studies

TL;DR Mayew and Venkatachalam ran the audio of company earnings calls through voice-analysis software and scored how much positive and negative emotion the chief executive and finance chief showed. They report that these vocal cues, especially while executives were being questioned by analysts, carried information about the company’s financial future that neither the numbers nor the words contained. The software itself is disputed: a phonetician has argued it cannot measure emotion at all.

  • Paper: Mayew, W. J., & Venkatachalam, M. (2012). The power of voice: Managerial affective states and future firm performance. The Journal of Finance, 67(1), 1–43.
  • Material: Audio recordings of quarterly earnings conference calls, split by speaker so that each executive could be analysed alone.
  • Measure: Layered Voice Analysis (LVA), a commercial program from the company Nemesysco, customised for the study, scoring each executive’s positive and negative affect against their own calm baseline.
  • Compared with: The companies’ later earnings, analysts’ forecasts and recommendation changes, and share prices.
  • Access: Subscription-only. The findings here come from the published abstract; how the measures were built comes from the authors’ 2008 working version, whose results differ in detail from the published paper.

Four times a year, the managers of public companies explain their results on a conference call and then take questions from analysts. Investors study every word of those calls. William Mayew and Mohan Venkatachalam, at Duke University, asked whether the voice carries something the words do not: whether how executives sound, particularly under questioning, says anything about how the company will do.1

The idea is simple. Executives know more about their company than anyone listening. They choose their words carefully, but emotion is harder to script, and a finance chief who is uneasy about next quarter may sound it before the numbers show it.

Turning a call into a score

The practical problem is measuring emotion in thousands of minutes of speech. The authors chose not to build their own measure from acoustic features, because, as they wrote, the research literature did not agree on which model to use. Instead they used a commercial program, Layered Voice Analysis, adapted by its maker so that it would export its readings.

The program reports many indicators. The authors used two. An “emotion level” reading, which the software’s makers present as excitement, stood for positive affect; a “cognition level” reading, presented as cognitive conflict or doubt, stood for negative affect. Each call’s score was simply the share of readings above the level the makers call critical.

What the voices predicted

What the study reports, in the published paper and the early working version

Published paper (2012)Working version (2008)
Both positive and negative affect were informative, when executives were being questioned by analystsMore negative affect: less likely to meet or beat earnings expectations in each of the next three quarters
Did not incorporate the information when forecasting near-term earningsNegative affect was not linked to revisions of next quarter’s forecast
Changes reflected positive affect but not negativeNot reported
Not stated in the abstractOver the next 180 trading days, returns were lower after more negative affect: the market caught up slowly

The claim is not that a voice predicts share prices on its own. It is that the voice adds something: in the authors’ words, information incremental to the reported numbers and to the soft information carried by what executives say. And the market seems to have been slow to use it, which is why returns could drift after the call.

The detail that the signal was clearest under questioning fits the story. Prepared remarks are rehearsed; answers to an analyst’s unexpected question are not.

Has it held up?

The idea has been influential in finance. The measurement has not been accepted outside it.

In 2012 Francisco Lacerda, a phonetician at Stockholm University, published a detailed critique arguing that the study’s results were void because the software cannot measure emotion.2 His argument rests on how the program, as described in its own patent, handles sound.

The phonetician’s objections, in numbers

What the software does, per the critiqueWhy it matters
Samples the voice 11,025 times a second at 8 bits, then divides the amplitude by 3, leaving about 85 levelsFine detail in the waveform is lost before anything is measured
Counts “thorns” and “plateaus”: tiny peaks and flat runs between three neighbouring samplesThese depend on the particular sounds, the microphone and the room, not on the speaker’s feelings
Looks at chunks of about 272 microseconds, treated as unrelatedEmotion is carried by pitch and rhythm over whole phrases; one pitch period of a 200 Hz voice spans about 18 such chunks
The algorithm is proprietaryIts errors cannot be checked or reproduced by other researchers

The authors had anticipated part of this. According to the critique, which quotes the published paper, they ran exploratory tests showing the software’s readings were related to measurable acoustic features, and concluded that this rejected the claim that the software extracts nothing relevant. The critique’s position is that such correlations are what a program would produce if its readings tracked loudness, microphones and background noise, and do not show that it measures emotion.

What survives, even on the sceptical reading, is narrower than the title. Something in the recordings of these calls, as scored by this program, went with later results. Whether that something is the executives’ emotion, as the paper argues, or a property of how they spoke and were recorded, is the question the critique raises and the program’s secrecy makes hard to settle.

What this study cannot tell you

  • Whether emotion was measured at all. Everything depends on a commercial program whose method is secret and whose validity is disputed by specialists in speech.
  • How large the effects are. We could not read the published paper, and the early working version’s numbers do not carry over.
  • Whether listeners can hear it. The measure was produced by software; the working version notes that the cues may be subtle or even inaudible, so it says little about what analysts or investors perceive.
  • Why the voice changed. A call that goes badly, a hostile question or a poor line could all move the scores.
  • Whether it lasts. Any predictable drift in share prices tends to shrink once it is published and traded on.

What a speaker can take from it

Even setting the software aside, the study rests on an idea most speakers will recognise: people listen for how you sound, not just what you say, and the unrehearsed moments give the most away. Executives were most revealing not in their prepared remarks but when an analyst asked something they had not scripted.

For anyone who speaks under questioning, the lesson is to prepare for the questions as seriously as for the talk. The prepared part is where you control the words; the questions are where the audience decides what you really think.

This article summarises a subscription-only paper. Its findings are taken from the published abstract; the method and the working-version results from the authors’ 2008 working paper, posted by New York University; the objections from Lacerda’s published critique. No figure is drawn from the published paper’s tables, which we could not access.

Measured

  • Readings from a commercial voice-analysis program on earnings-call audio, separately for chief executives and finance chiefs
  • Later earnings, analysts’ forecasts and recommendation changes, and share prices

Inferred, not measured

  • That the program’s readings are emotion: asserted by its makers and the authors, disputed by the critique
  • That analysts and markets under-used the information: inferred from how forecasts, recommendations and prices moved

References

  1. Mayew, W. J., & Venkatachalam, M. (2012). The power of voice: managerial affective states and future firm performance. The Journal of Finance, 67(1), 1–43. https://doi.org/10.1111/j.1540-6261.2011.01705.x
  2. Lacerda, F. (2012). Money talks: the power of voice. A critical review of Mayew and Venkatachalam’s The power of voice: managerial affective states and future firm performance. PERILUS 2012, Department of Linguistics, Stockholm University. https://www.diva-portal.org/smash/get/diva2:509721/FULLTEXT01.pdf