Humans read emotion in voices constantly and mostly without effort. We hear irritation in a clipped "haan, theek hai", warmth in a greeting, anxiety in a rushed explanation. It is natural to want machines to do the same, especially in customer service, where a frustrated caller who is not noticed may become a lost customer.
Speech emotion recognition (SER) is a well-established research field with decades of work behind it. It has also attracted marketing claims that go well beyond what the evidence supports. This post explains how SER works, why some parts of the problem are much harder than others, what changes on a phone line and in Indian languages, and where ethical and legal lines are being drawn.
What the Voice Carries Beyond Words
Speech carries two broad streams of information: the linguistic content (what is said) and paralinguistic information (how it is said). Paralinguistic cues are the raw material of emotion recognition from audio. The most studied include:
- Pitch (fundamental frequency): its average level, range and variability. Heightened activation often raises pitch and widens its range.
- Energy (loudness): overall intensity and how it changes. Anger and excitement tend to be louder; sadness tends to be quieter.
- Speaking rate and rhythm: syllables per second, pause frequency and duration.
- Voice quality: breathiness, tension, creakiness, harshness and jitter or shimmer (cycle-to-cycle variation in pitch and amplitude).
- Spectral characteristics: how energy is distributed across frequencies, which is affected by vocal effort and articulation.
Classical SER systems computed hand-engineered feature sets from these cues and fed them to statistical classifiers. Standardised feature sets were developed in the research community to make results comparable. Modern systems increasingly learn representations directly from audio using neural networks, often starting from models pretrained on large amounts of unlabelled speech, and many combine acoustic cues with the transcribed words.
Categorical and Dimensional Models of Emotion
Before a system can recognise emotion, someone has to decide what counts as an emotion. Two traditions dominate.
Categorical Models
Categorical approaches assign discrete labels such as anger, happiness, sadness, fear, surprise, disgust and neutral. This tradition is associated with Paul Ekman's work on basic emotions. Categories are intuitive and easy to act on, but real feelings are often blended ("frustrated but polite"), and different datasets use different label sets, which makes models hard to compare or transfer.
Dimensional Models
Dimensional approaches describe emotion as a point in a continuous space. James Russell's circumplex model, published in 1980, placed emotions on two axes, and a common extension uses three:
- Valence: how positive or negative the state is.
- Arousal (activation): how calm or energised it is.
- Dominance: how in control or submissive the person feels.
Dimensional labels capture subtlety and mixed states better, but they are harder for human annotators to rate consistently and less direct to act on.
Acted Versus Natural Speech
Much foundational SER research relied on acted emotion, where actors portray specified emotions. Acted data is easy to label, but portrayals tend to be clearer and more exaggerated than everyday emotion.
A widely used dataset, IEMOCAP (Interactive Emotional Dyadic Motion Capture), developed at the University of Southern California and described by Busso and colleagues in 2008, sits partway along this spectrum. It contains recordings of actors in pairs, performing both scripted scenes and improvised scenarios intended to elicit natural reactions, annotated with both categorical and dimensional labels. Later efforts have collected more naturalistic speech, for example from podcasts or real conversations.
The consistent lesson across the field is that performance drops on natural speech. Real emotions are usually subtle, mixed and masked by social norms. A customer who is annoyed often sounds perfectly polite. Human annotators themselves frequently disagree about the emotion in natural recordings, which sets a practical ceiling on how well any system trained on their labels can perform.
Why Arousal Is Easier Than Valence
Research has repeatedly found that arousal is predicted from acoustics considerably better than valence. The reason is physiological and acoustic. High activation, whether from anger, excitement or fear, tends to change the voice in similar, measurable ways: higher pitch, more energy, faster speech, tenser voice quality.
Valence is different. Hot anger and joyful excitement can both be loud, fast and high-pitched. Distinguishing them from sound alone is difficult; the words and context often carry more information than the acoustics. This is one reason modern systems increasingly combine audio with text. It also means that a system confidently labelling a loud caller as "angry" may simply be detecting energy.
Language, Culture and Individual Variation
Emotional expression in voice is shaped by language, culture, social setting and individual habit.
- Prosody serves grammar too. Pitch and stress patterns carry linguistic meaning, and those patterns differ between languages. A pitch contour that sounds emphatic in one language may be ordinary in another.
- Display norms vary. Cultures and contexts differ in how openly emotion is expressed, particularly towards strangers or service providers.
- Code-mixing and multilingual speakers. In India, a caller may shift between Hindi, Gujarati and English, and prosodic habits can shift with the language.
- Individual baselines differ. Some people naturally speak loudly and quickly. Without knowing a speaker's baseline, a system may mistake their normal voice for agitation.
Most widely used SER datasets are in English and a small number of other languages. Evidence for how well models transfer to Indian languages and Indian conversational styles is limited, so claims of cross-lingual accuracy deserve scepticism unless they are backed by evaluation on relevant data.
The broader scientific debate is also relevant. A 2019 review by Lisa Feldman Barrett and colleagues, focused on facial expressions, argued that emotional states cannot be reliably inferred from facial movements alone because the same expression can mean different things in different contexts. Similar caution applies to voice: acoustic patterns correlate with emotional states, but they are not a reliable readout of what someone feels.
The Telephony Effect
Phone calls are a harsh environment for paralinguistic analysis. Traditional narrowband telephony carries roughly 300 to 3,400 Hz, discarding low and high frequencies. Many calls also pass through lossy codecs, packet loss concealment, automatic gain control and noise suppression.
These factors affect exactly the cues emotion recognition depends on:
- Energy is distorted by gain control and varies with how close the phone is held.
- Voice quality and spectral cues lose information above the band limit.
- Background noise, such as traffic, a television or a busy shop, adds energy that has nothing to do with the speaker.
- Codec artefacts can resemble or mask cues like breathiness or roughness.
Models trained on clean, wideband studio recordings typically degrade on telephone audio. Training or at least evaluating on realistic telephony data is essential before trusting any output.
Tone Is Not Satisfaction
There is a crucial gap between detecting a vocal pattern and understanding a customer's state. A system might accurately detect high arousal and a tense voice. That does not establish that the customer is unhappy with the business. The caller might be in a hurry, standing in a noisy market, stressed about something unrelated, or simply a naturally emphatic speaker.
Conversely, a calm and polite caller may be about to cancel their subscription. Many dissatisfied customers never raise their voice. What they say ("yeh teesri baar ho raha hai", this is happening for the third time) is often a far better signal than how they say it.
Useful customer insight usually combines several sources: the words spoken, the conversation's outcome, repeated contacts, and explicit feedback. Acoustic emotion estimates can contribute, but they should not stand in for these.
Ethics and Regulation
Inferring emotion about people raises serious concerns: it is intrusive, error-prone and potentially discriminatory if models behave differently across accents, languages, ages or genders.
Regulators have begun to respond. The European Union's AI Act (Regulation (EU) 2024/1689) prohibits placing on the market or using AI systems to infer the emotions of natural persons in the areas of workplace and education institutions, except where the system is intended for medical or safety reasons. According to the Act's timeline, the prohibitions began to apply in February 2025. Emotion recognition systems outside those prohibited contexts are not banned outright, but the Act classifies certain biometric emotion recognition uses as high-risk and imposes transparency obligations, such as informing people exposed to such systems. The precise scope, including definitions and guidance from the European Commission, is detailed, and organisations operating in or serving the EU should seek legal advice.
A practical implication for customer service is worth noting: analysing the emotions of employees, such as call centre agents, sits much closer to the prohibited workplace use case than analysing customers does. Even outside the EU, these rules are a useful reference for what regulators consider unacceptable.
Realistic Uses and Overclaims
Used with humility, speech emotion signals can be genuinely helpful.
Reasonable Applications
- Escalation hints: flagging calls where sustained high arousal coincides with complaint language, so a human can step in sooner.
- Quality assurance sampling: prioritising which calls supervisors review, rather than sampling at random.
- Aggregate trends: noticing that calls about a particular issue are consistently more tense, which points to a process problem.
- Conversation design feedback: identifying moments in automated flows where callers commonly become frustrated.
Claims to Treat With Caution
- Detecting lies, intent or trustworthiness from voice.
- Scoring individual employees on their emotional state.
- Precise, fine-grained emotion labels from short telephone clips.
- Accuracy figures quoted without stating the dataset, language, recording conditions and whether speech was acted or natural.
Looking Forward
Speech emotion recognition is a real capability with real limits. It is reasonably good at noticing that something in a voice has changed, especially in intensity, and much weaker at knowing what a person actually feels or why. Advances in self-supervised speech models and multimodal systems that combine audio and text are improving results, but the fundamental ambiguity of emotional expression will not disappear.
For businesses, the responsible approach is clear: treat inferred emotion as a signal that prompts human attention, never as a verdict about a person; evaluate on your own language and telephony conditions; be transparent with callers; and stay well away from uses that regulators and ethicists have identified as harmful.
Frequently Asked Questions
What is speech emotion recognition?
Speech emotion recognition is the task of inferring emotional states from spoken audio, using cues such as pitch, loudness, speaking rate and voice quality, sometimes combined with the words spoken. Systems output either emotion categories like anger or sadness, or continuous dimensions such as arousal and valence.
How accurate is emotion recognition from voice?
Accuracy depends heavily on the data, the emotions being distinguished and the recording conditions. Systems tend to perform much better on acted speech than on natural conversations, and better at detecting activation or intensity than at telling positive from negative feelings.
Is emotion recognition allowed under the EU AI Act?
The EU AI Act prohibits AI systems that infer emotions of people in workplaces and educational institutions, with exceptions for medical or safety reasons, and places other emotion recognition systems under additional obligations. The details and their interpretation are complex, so organisations operating in the EU should seek legal advice.
What are responsible uses of emotion detection in customer calls?
Responsible uses treat acoustic signals as hints rather than verdicts, for example prioritising calls for human review, flagging possible escalation needs, or sampling calls for quality assurance. Decisions about individuals should not rest on inferred emotion alone.