AI-Powered Key Takeaways
Audio can be technically functional and still sound unpleasant. A video stream can load successfully and still look blurry, blocky, or unstable. Metrics such as latency, packet loss, bitrate, and frame rate help explain what is happening technically, but they do not always answer a more basic question: How good does the experience actually seem to the user?
That is where the mean opinion score comes in.
Mean Opinion Score, commonly shortened to MOS, is a way of expressing perceived media quality as a numerical score. It originated with human listeners rating communication quality and has since expanded into objective and model-based approaches for assessing voice, video, and audiovisual experiences.
This guide explains what MOS measures, how a Mean Opinion Score calculation works, how MOS is applied to voice and video, what affects the score, and how teams can use it alongside other performance metrics.
What Is MOS? (Quick Definition)
Mean Opinion Score (MOS) is a numerical representation of perceived audio, video, or audiovisual quality. The traditional MOS scale runs from 1 to 5, with higher scores representing better perceived quality.
The commonly used Absolute Category Rating scale is:
The concept is simple. A group of people experiences the same audio or video sample, each person assigns a rating, and the ratings are averaged. That average is the mean opinion score.
ITU-T defines MOS terminology for audio, video, and audiovisual quality and distinguishes between scores obtained from subjective tests, objective models, and network planning models.
That distinction matters. Two values may both be called MOS, but they may have been produced in very different ways.
A MOS of 4.1 from a controlled human listening test, for example, should not automatically be treated as equivalent to a 4.1 generated by a network-based prediction model. ITU-T guidance recommends reporting enough information about the test and methodology for the value to be interpreted correctly.
Also read : Harnessing AI to Track and Optimize Video Quality
The History and Standards Behind MOS
MOS has its roots in telecommunications, where engineers needed a repeatable way to evaluate how people perceived telephone transmission quality.
Rather than judging a connection only through electrical or network measurements, subjective testing introduced the listener's experience into the assessment. Participants could listen to speech samples under controlled conditions and assign quality ratings, producing a score that reflected human perception.
ITU-T Recommendation P.800, approved in its current numbering in 1996 after the earlier P.80 recommendation was renumbered, describes methods for conducting subjective evaluations of transmission quality.
Several related ITU-T recommendations now help define how MOS should be used:
- ITU-T P.800: Methods for subjective determination of transmission quality.
- ITU-T P.800.1: Defines terminology for MOS across audio, video, and audiovisual quality and distinguishes subjective, objective, and estimated scores.
- ITU-T P.800.2: Provides guidance for interpreting and reporting MOS values.
- ITU-T P.863: Defines an objective method for predicting perceptual speech quality.
- ITU-T G.107: Defines the E-model used for transmission planning and estimating voice quality from impairment factors.
- ITU-T P.1203: Covers model-based quality assessment for progressive download and adaptive audiovisual streaming.
MOS has therefore evolved beyond its original human listening-test format. The core idea remains the same, though: represent perceived quality in a form that can be measured and compared.
Also Read : How to Measure Video Quality Effectively
How Is MOS Measured/Calculated?
There is no single Mean Opinion Score calculation that applies to every MOS value. The calculation depends on how the score is obtained.
1. Subjective MOS testing
The traditional method uses human participants.
Participants listen to or view the same test material and assign ratings using a predefined scale. With a standard 1-to-5 Absolute Category Rating scale, each participant selects a score from 1 to 5.
The MOS is then the arithmetic mean:
MOS = Sum of all individual opinion scores ÷ Number of participants
For example, suppose five listeners give a voice sample the following ratings:
4, 4, 5, 3, 4
The Mean Opinion Score calculation would be:
(4 + 4 + 5 + 3 + 4) ÷ 5 = 4.0
The resulting MOS is 4.0.
Subjective testing has an obvious advantage: it measures human perception directly.
It also has practical limitations. Recruiting participants, controlling test conditions, presenting samples consistently, and collecting enough responses takes time. That makes continuous testing at scale difficult.
2. Objective perceptual models
Objective models use algorithms designed to predict how human participants are likely to rate media quality.
For voice, one example is POLQA, standardized under ITU-T P.863. It estimates perceptual listening quality using a reference speech signal and the corresponding degraded signal. ITU-T describes P.863 as an objective method for predicting listening speech quality.
Another name teams may encounter is PESQ, which was standardized under ITU-T P.862. However, P.862 and its related recommendations were withdrawn by ITU-T on January 5, 2024, with ITU-T directing users to the P.863 family instead.
3. Model-based or estimated MOS
MOS can also be estimated from technical parameters rather than directly analyzing what a listener hears.
For VoIP, the E-model defined by ITU-T G.107 uses transmission impairment factors to calculate an R-factor. That rating can then be mapped to an estimated MOS value.
This makes estimated MOS useful for network planning and monitoring because quality can be evaluated without assembling a listening panel for every call.
The important point is that MOS is the resulting quality score, while the method used to produce that score may differ considerably.
MOS for Voice & VoIP Calls
Voice communication remains one of the most common applications of mean opinion score MOS testing.
During a traditional subjective test, listeners hear speech transmitted through the system under evaluation and rate its perceived quality. In production environments, however, voice quality is more commonly estimated using objective or network-based models.
Several factors can influence perceived VoIP quality.
1. Packet loss
Voice traffic is time-sensitive. If packets are lost in transit, portions of the audio may be missing or reconstructed through packet-loss concealment.
Small amounts may be difficult for a listener to notice. As loss increases, however, speech can become distorted, broken, or difficult to understand.
2. Jitter
Jitter describes variation in packet arrival timing.
Real-time audio expects packets to arrive at reasonably predictable intervals. Large variations can force the receiving system to buffer packets or discard packets that arrive too late for playback.
3. Latency
Latency affects how quickly speech travels between participants.
A call may still sound clear even with additional latency, but long delays interfere with normal conversational timing. People may begin speaking over one another or experience noticeable pauses between speaking and receiving a response.
4. Codec choice
Voice codecs compress and encode speech differently. Codec characteristics, bitrate, bandwidth, packet-loss handling, and other implementation details can affect perceived quality.
5. Audio capture and playback conditions
Not every quality problem comes from the network. Microphones, speakers, acoustic environments, background noise, echo, device processing, and application behavior can also influence what the listener actually hears.
This is why MOS should not be treated as a substitute for the supporting measurements. A low score tells you that perceived quality has degraded. Packet loss, jitter, latency, codec information, and other diagnostics help determine why.
Also read - How Network Latency Impacts Mobile Gaming Experience
MOS for Video & Streaming Quality
MOS is also used to describe perceived video and audiovisual quality.
The basic goal remains the same: turn the viewer's perception into a score that can be tracked and compared. The factors affecting video perception, however, are different from those affecting a voice call.
Video quality can be influenced by:
- Resolution and image detail
- Compression artifacts
- Blurriness and blockiness
- Frame rate and frame consistency
- Playback stalls and buffering
- Initial playback delay
- Changes in bitrate or resolution
- Audio quality and audio-video synchronization
- Device and display characteristics
- Network conditions
For streaming applications, the overall experience is especially important.
A video can look excellent while it is playing but still provide a poor experience if playback repeatedly freezes. Likewise, a stream with no buffering may still look unpleasant if aggressive compression creates visible artifacts.
This is one reason audiovisual quality models consider more than image quality alone.
For example, ITU-T P.1203 defines a model for adaptive audiovisual streaming that combines short-term audio and video quality estimates with information such as initial loading delay and playback stalling to estimate overall session quality.
MOS gives teams a useful perceptual layer, but its value increases when it is analyzed alongside the underlying media, playback, device, and network measurements.
Where Is MOS Used?
MOS is useful wherever teams need to understand perceived media quality rather than relying only on infrastructure or application health.
Common use cases include:
1. VoIP and telephony
Telecommunications providers and communications platforms can use MOS or estimated MOS to track speech quality across calls, codecs, network paths, and service conditions.
2. Video conferencing
Video conferencing combines real-time audio, video, networking, and device performance. MOS-based measurements can help teams assess whether participants receive an acceptable media experience under different conditions.
3. OTT and video streaming
Streaming teams can evaluate how content appears across devices, network conditions, encoding profiles, and playback scenarios.
4. Live streaming
Live video introduces conditions where a clean reference video may not always be available during testing. Reference-free perceptual quality models can be particularly useful in these situations.
5. Gaming and interactive media
Cloud gaming, game streaming, and other interactive media experiences depend on consistent visual quality as conditions change. Perceptual quality scoring can complement frame rate, latency, rendering, and network measurements.
6. Codec and media pipeline evaluation
Teams can compare encoding configurations, compression levels, transcoding behavior, and media-processing changes to understand how they affect perceived quality.
7. Device and network testing
The same media experience can behave differently depending on the device, display, network, geographic path, and available bandwidth. Testing across representative environments helps reveal those differences.
Also read - OTT Platform Testing: A Guide to a Seamless Streaming Experience
How to Improve a Low MOS Score
A low MOS is a signal that quality needs attention, not a diagnosis by itself.
The right response is to determine which part of the experience is contributing to the score.
1. Look at the supporting metrics
Start by examining technical measurements recorded during the same session.
For voice, that may include:
- Packet loss
- Jitter
- Latency
- Codec
- Bitrate
- Audio levels
For video, examine:
- Buffering
- Frame rate drops
- Resolution changes
- Blurriness
- Blockiness
- Bitrate
- Playback startup time
- Audio-video synchronization
The goal is to identify whether the change in MOS coincides with a measurable change elsewhere.
2. Test network conditions
Media applications should be evaluated under the types of connections users actually encounter.
Compare performance across Wi-Fi, cellular networks, different bandwidth conditions, and geographic locations where relevant. Network shaping can also help isolate how specific latency, bandwidth, or packet-loss conditions affect the experience.
3. Review codec and encoding configurations
For voice, confirm that the chosen codec and configuration are appropriate for the expected network environment.
For video, review bitrate ladders, compression settings, resolution, frame rate, and adaptive bitrate behavior. Increasing bitrate alone is not always the answer. The objective is to find a configuration that maintains acceptable perceptual quality while working within realistic delivery constraints.
4. Investigate playback behavior
For streaming applications, check whether poor MOS regions coincide with buffering, resolution changes, dropped frames, or other playback events.
This can help distinguish a consistently poor encode from a quality problem that appears only during specific parts of the session.
5. Test across representative devices
A result from one device does not prove that every user receives the same experience.
Differences in screen characteristics, processing capability, operating system behavior, media decoding, and application implementation can affect playback. Repeating the same journey across representative devices makes the result more meaningful.
6. Compare like with like
Avoid comparing MOS numbers without considering how they were produced.
ITU-T guidance notes that values from separate subjective experiments are not necessarily directly comparable unless the experiments were designed for that purpose. Test methodology and context should therefore remain consistent when MOS is used for benchmarking or regression analysis.
Also read - Benchmarking Video Performance in 5G Networks for Streaming Apps
MOS vs. Other Quality Metrics
MOS is most useful when it complements other measurements rather than replacing them.
There is an important pattern here.
MOS answers a perception question. Most supporting metrics answer technical questions.
If MOS falls at the same point that packet loss increases, frames begin dropping, or blockiness rises, the combination tells a much more useful story than either measurement alone.
Best Practices for Testing MOS with HeadSpin
HeadSpin's MOS capability is focused on perceived video quality through Video Mean Opinion Score (VMOS).
HeadSpin VMOS provides a 1-to-5 perceptual video quality score based on a model trained using human-labeled data. It can evaluate video without requiring a source reference, which makes it applicable to scenarios such as live streams and other dynamic video content.
When using VMOS with HeadSpin:
- Test on representative real devices and networks. Evaluate the media experience under the device and connectivity conditions relevant to your users.
- Read VMOS alongside other video KPIs. HeadSpin can capture measurements including blurriness, blockiness, frame rate drops, and buffering time, helping teams investigate what was happening when perceived quality changed.
- Use consistent scenarios when comparing results. Keep content, test flow, device conditions, and other variables controlled where possible so changes in VMOS are easier to interpret.
When a reference video is available, HeadSpin also supports VMAF-based full-reference video quality analysis, giving teams another way to evaluate the output alongside reference-free VMOS and other media KPIs.
Conclusion
Mean Opinion Score gives teams something that raw network and media metrics cannot provide on their own: a numerical view of perceived quality.
The original approach was straightforward. People listened to or watched content, rated what they experienced, and those ratings were averaged. Modern systems can also estimate perceptual quality through objective models and network-based calculations, allowing MOS-style measurements to be used at much greater scale.
But a MOS number should never be read without context.
First understand how the score was produced. Then examine the measurements behind it. A falling score becomes far more useful when you can connect it to packet loss, buffering, compression artifacts, frame drops, or another specific part of the media experience.
That combination of perceptual and technical measurement is what turns MOS from a quality rating into a useful testing signal.
FAQs
Q1. What affects MOS in VoIP calls?
Ans: VoIP quality can be affected by packet loss, jitter, latency, codec behavior, audio equipment, acoustic conditions, and application processing. These measurements should be examined alongside MOS to determine why perceived call quality changed.
Q2. What affects video MOS?
Ans: Video MOS can be influenced by compression artifacts, resolution, blurriness, blockiness, frame consistency, buffering, bitrate changes, audio-video synchronization, and other factors that affect the viewing experience.
Q3. Is MOS the same as VMAF?
Ans: No. MOS represents perceived quality, traditionally through human ratings or through models designed to estimate perception. VMAF is a full-reference video quality metric that evaluates a processed video relative to a reference source.
Both can be useful, but they answer the quality question differently.
Q4. Can MOS be used for video streaming?
Ans: Yes. MOS-based approaches can be used to evaluate perceived video and audiovisual quality. ITU-T P.1203, for example, defines model-based quality assessment for adaptive audiovisual streaming and incorporates video, audio, initial loading delay, and playback stalling into its quality model.
.png)







.png)















-1280X720-Final-2.jpg)








