Saturday, March 3, 2012

MOS Results are Good but Customers Still Complain of Poor Speech Quality

When monitoring the speech packets of a VoIP network, i.e. the RTP stream, protocol analyzers or VoIP service monitoring systems may analyze the packets and calculate a Mean opinion score (MOS) which describes the quality of the speech. If customers call into your call center or NOC and complain about speech quality this is the first place you would go to analyze the problem. Finding the bad call the customer is referring to will take a few seconds with a good voice service assurance system Sometimes when you find the call, the MOS maybe 4.0 or greater indicating good speech quality.

How can we explain this discrepancy?

The method used to calculate MOS by a monitoring system which needs to monitor thousands of active calls and produce voice quality measurements or QoS of all of them is R-factor based on the E-model (ITU-T G.107). R-factor takes into account the codecs used by the endpoints (VoIP Phones) in the call. But most of the calculation is derived from what happens to the packet stream as it transits an IP packet network. So R-factor requires measurements for packet loss and jitter of the packets in the stream of VoIP (RTP) packets

How is packet loss and jitter detected in an RTP stream? The RTP packets contain a sequence number and a time stamp and from this packet loss and jitter can be calculated. It is however, important to understand that these measurements can only be made on the received RTP streams i.e. packet loss and jitter is only measured up to that point in the network where the monitoring system probe is connected.

RTP streams leaving your network going to your customer’s premises will not be measured unless you have a probe on the customer’s premises or their VoIP phones support RTCP. A sophisticated good voice service assurance system will also read the RTCP packets coming back from the customer’s probes and so that leg of the call can be included in the R factor measurement. Some monitoring systems will show you in which leg of the call the packet impairments are added.

If call quality is bad but MOS is good, perhaps you are not monitoring all legs of the call.

Another reason may account for this discrepancy. MOS or Perceptual Speech Quality does not include echo or overall end-to-end delay. So if the customer is experiencing echo, this will not degrade the MOS values. Similarly, if there is delay in the network (as opposed to jitter, which is short-term varying delay), this will not impact the MOS value. Network Delay in VoIP causes cold or stinted conversation, so this this could be a reason for your customer to call you.

If customers complain of poor speech quality but the MOS results are good, use your voice service assurance system to record the calls from this customers (with disclaimers and forwarding necessary of course) and listen to a sample for echo or delay.

Saturday, February 25, 2012

Methods to Objectively Evaluate Speech Quality


This post gives an overview of methods to objectively evaluate perceptual quality of speech. The science of perceptual speech quality measurements assessment has been progressing for the past two decades. Current state-of-the-art for full referenced measurements uses the third iteration of the standard. Good-quality implementations of the standards of the algorithms contained within the standards will deliver the 97% correlation with the subjective tests database depending on implantation of the test set up.

Commercial test solutions exist to measure speech from acoustic, electronic analog and digital interfaces.

Mean Opinion Score [MOS] is a scale used in voice telecommunications predicting the perceived quality of a speech sample. MOS describes speech clarity or intelligibility. The measure does not measure delay across a network or Echo. The scales runs from 1 to 5, a score of '1' is bad and '5' is excellent. MOS [Mean Opinion Score] test sessions comprise 15 to 25 people listening to speech files of good quality and of poor quality with impairments and scoring them subjectively. This subjective test process is specified in the standard ITU-T P.800.

 
Mean opinion score (MOS)
MOS
Quality
Impairment
5
ExcellentImperceptible
4
GoodPerceptible but not annoying
3
FairSlightly annoying
2
PoorAnnoying
1
BadVery annoying

 
MOS is valuable for the characterization of any device where the voice is compressed and transmitted over networks. Such devices may be handsets, mobile phones and networks such as packet-based VoIP networks or wireless networks. Networks employing state-of-the-art codecs are optimized for the compression of voice and treat tomes such as DTMF tones in a different way from the voice. Therefore to the characterize quality of the network or the System under Test, real speech files need to be transmitted. Frequency response, levels, Echo and delay must be measured, but in addition to the perceived speech quality.

Subjective testing is obviously time-consuming and expensive. So algorithms have been developed to allow a computer to reference the pure incidented speech file and compare it with the received degraded file and calculate the MOS with a high degree of correlation to subjective MOS. The first such objective speech quality measurements were standardized in 1998 with the PSQM algorithm P.861 Objective quality measurement of telephone-band (300-3400 Hz) speech codecs. This was followed by ITU-T Recommendation P.862: Perceptual evaluation of speech quality (PESQ) which has gained widespread worldwide usage as a reliable method for characterizing most narrowband telephony systems.


The following are examples of Mean Opinion Scores for one implementation of different codecs:

Codec
Data rate
[kbit/s]
Mean opinion score
(MOS)
G.711 (ISDN)644.1
iLBC15.24.14
AMR12.24.14
G.72983.92
G.723.16.33.9
GSM
EFR
12.23.8
G.726 ADPCM323.85
G.729a83.7
G.723.15.33.65
G.728163.61
GSM
FR
12.23.5

 

 

 

 




 

 







Wideband telephony networks are expected to improve the user experience including the intelligibility of voice conversations over highly compressed codecs as used in both packet and wireless networks. Hopefully the tardy phrase "can you hear me now" will become a less frequent part of our vocabulary. 

High Definition or Wideband Telephony is just now coming into common usage. G.722 is the Wideband Telephony codec for VoIP and WB-AMR is now being tested for wireless networks. However, they do need to be tested to tune codec implementations, packet loss concealment algorithms and performance in areas of poor coverage or high congestion.

Here are the analog frequency definitions for the different forms of telecommunications:



 

 

 

 

  

  







PESQ was never designed to address wideband networks. In addition, vendors of time warping code codecs [e.g. EVRC] and Skype and iLBC were not content that PESQ accurately measured the full quality of their codecs.

 
In 2006 ITU-T commenced work on a new standard to address the limitations of PESQ and during 2011 the POLQA standard, Recommendation P.863, was published.

 

 

 


 




















Speech uses the new POLQA speech quality metric for objective protection of MOS. The old PESQ algorithm (ITU-T P.862) has been used for narrowband telephony since it was approved in 2000. PESQ was not designed for Wideband Telephony and also did not represent well the speech quality of time warping codecs. POLQA addresses all these short comings and provides a scale that goes all the way to 24kHz audio.

It is desirable to use the same scale so that laboratories can compare new results for wideband telephony with their old PESQ database. However, the question of human expectation comes into play because all these objective measurements performed by computers must correlate or predict subjective experience. If you watch a video on your smart phone, you might consider the picture quality as being good. Your expectations are put in the context of the small screen and the convenience of the video being played on a handheld smartphone. If you would give you the same video on your brand-new expensive high-definition 1080P TV, you would be very disappointed even if the pixel resolution had been scaled to the 62 inches screen size. Your expectation of quality is tempered to the format in which you are viewing it.

Similarly with speech and audio. If you were to participate in a MOS test and invited into a studio where there were high fidelity speakers, orchestral classical music playing and told and asked to rate the quality of the High Definition speech you are about to hear, your expectations would be set high and you'd be more critical. You would score the audio lower than if you had been asked to rate the speech quality of your most recent cellular phone call.

POLQA offers two scales, the narrowband scale and the super wideband scale. Super wideband telephony reaches 14 kHz analog audio frequency. The narrowband focus scale maps directly onto the old desk scale and exploits the higher scores not given by test participants in narrowband tests.

• NB: Maximum MOS value 4.25

• WB: Maximum MOS value 4.5

• SWB: Maximum MOS value 4.75

So a score of 4.5, on the narrowband POLQA scale is experimentally the best value you will ever obtain with wideband telephony equipment. For POLQA, the maximum MOS value in tests is 4.75 .

In future years, the industry will migrate exclusively to using the super wideband POLQA scale as soon as users' expectations always expect high-definition or hi-fi quality to the communications audio.

POLQA SWB
POLQA NB
14kHz 16 bit Linear
4.75
7kHz 16 bit Linear
4.5
AMR - WB
4
3.4KHz 16 bit Linear
3.8
4.5
G.711
3.7
4.3
EFR/AMR-FR 12.2kbps
3.6
4.1
EVRC 9.5 kbps
3.4
3.9
EVRC-B 9.5 kbps
3.5
4
AMR-HR 7.95 kbps
3.4
3.8

 

Applications for Perceptual Speech Quality Measurements 

Perceptual speech quality measurements are used to make End-to-End measurements, for any network where voice codecs are used to compress the speech or where speech transmission systems or networks may introduce impairments, such as weak radio signals or multipath or packet loss or packet jitter.
Examples of systems where Perceptual Speech Quality Measurements are valuable:

  • Codec evaluation
  • Frame or packet concealment implementation
  • Headsets combining the digitization of sound
  • VoIP phone assessment, both softphone and dedicated VoIP phones
  • Mobile handsets
  • Digitizing radio & Intercom systems
  • VoIP networks
  • Wireless cellular networks
  • speech enhancement and noise reduction systems
  • transcoders

A new important feature added to POLQA is its ability to measure the improvement of speech quality for speech enhancement and noise reduction systems.

The picture shows the iLBC codec measuring 4.21 and the narrowband
POLQA scale.


























In over 2 decades where these tests have taken place, no statistically significant number of participants ever scored any speech recording as being excellent or 5.0. The highest score typically obtained in any test was 4.54. So this measurement for iLBC of 4.21 is a good score for the codec.

For more information on making PESQ and POLQA measurements, ensure you contact only renowned and well-respected test vendors because the science of speech quality measurements requires expertise and experience in many different areas - audio, analog electronics as well as computing. It is easy to make a measurement but care is required to ensure that measurement is accurate correlates to human subjective experience and is put into the context of the environment, resolution, format etc.

one of the most admired vendors for
 speech quality metrics is Malden Electronics, available in USA through Teraquant Corporation – 
www.teraquant.com


Sunday, February 5, 2012

Some Key Facts About Communications Fraud

According to a recent study from the Communications Fraud Control Association (CFCA), a high number of Communication Service Providers are increasingly facing fraud attacks. Based on the results of the CFCA’s “2011 Global Fraud Loss Survey”, the annual Global Fraud Loss Estimate reportedly is $40.1 Billion (USD), whereas the real number is probably much higher.

The CFCA which is a not-for-profit global educational association with mission to combat communications fraud, describes the 5 most current fraud types as follows:
  • $4.96Billion–Compromised PBX/Voicemail systems
  • $4.32Billion–Subscription/Identity (ID)Theft
  • $3.84Billion–International Revenue Share Fraud (IRSF)
  • $2.88Billion–Bypass Fraud
  • $2.40Billion–Credit Card Fraud

However, there are many other existing fraud types such as Identity Theft, Social Engineering, Roaming Fraud, SS7 Manipulation just to name a few.

When asked how many fraud incidents the surveyed telecom operators handle per month, 42.6% said less than 50 incidents, 13% report to have had 51-100 incidents, 24.1% (101-500 incidents), 3.7% (501-1,000 incidents) and 16.7% (1,001+ incidents).

With the migration to Next Generation Networks, the risk of new fraud attack methods keeps on growing and the need to protect these networks becomes an obvious necessity for network operators. Current available solutions to fight telecom fraud might still be limited, but to prevent fraud attacks efficiently, you might need a solution that runs in real-time to detect fraud incidents instantly, with a powerful technology to prevent any fraud attack on IMS, LTE and VoIP networks.

Source: all provided numbers and facts are based on the CFCA’s “2011 Global Fraud Loss Survey”.

For more information about the Communications Fraud Control Association, please contact fraud@cfca.org or visit the official website: www.cfca.org


Wednesday, January 25, 2012

The One-Way Audio VoIP Dilemma

One-way audio is an annoying anomaly where one person can hear the other caller, but that caller can't hear them.  It is uncommon in non-VoIP telephony, but seems to occur frequently in systems that have just been deployed.

In traditional TDM voice transmissions, the circuits are reserved for two-way voice transmission, but in VoIP each voice stream is independent.  Should one of the streams be lost, deferred, or misdirected, the call the results on has one active stream in one direction.

The most common reason for one-way audio is the result of improper negotiation of the RTP voice stream by a server that is behind a firewall - an unaccommodating firewall.

In a future post we will talk about how to isolate the offending firewall and how to work with the system or firewall administrator to set NAT parameters appropriately so that both RTP streams can be negotiated successfully.

Thursday, January 5, 2012

Attacking VoIP Call-Quality Issues

Call-quality issues are a bit trickier with VoIP than with traditional telephony.

One of the main problems with VoIP are the negative effects of the delays in transmission of packets.  The latency is always inherent in converting a call to and from VoIP.  But no matter how quickly that conversion occurs, if there are several, the cumulative effective latency can have noticeable degradation in quality.  An end-to-end cumulative latency of great than 200 milliseconds (ms) causes human perceptible delay and results in conversations with participants that repeatedly interrupt each other.

Fluctuations in latency are also common on VoIP networks.  This is commonly referred to as jitter.  Jitter is in most cases caused by a leg in the call path that is sharing traffic with other applications or calls.  The most common  symptom of jitter is called "clipping."

A normal conversation sounds like this:
"Hi dear, shall I stop at the store to pick up some milk?"
but clipping produces:
"I ear, all op de ore ick om ilk"

In future posts we'll explore how to isolate latency and jitter to one or many legs of a VoIP call.

Thursday, December 15, 2011

The Essence of VoIP Issues

In the telephony world VoIP is a hybrid.

A VoIP solution usually has dedicated Internet connections, dedicated T1 voice circuits, and at two or or more points in a call a digital to analog conversion.

On top of these common telephony components, you add the complexities associated with IP packets:  latency, real time transport of packets, call control channels translations, and dial plan nuances.

One of the first questions to ask yourself when troubleshooting a VoIP call is whether or not a problem exhibits itself on all of your calls or just select ones.

In future posts we will begin to begin to isolate trouble to the various components of the call circuit and learn what clues to look for with each isolated issue or problem.

Friday, November 18, 2011

The True Customer Experience of a Voice Service

Customers perceive the quality of a voice service and the reliability of a voice service based on many different experiences. Such experiences may be good or bad, but very often customers notice only the bad experience – obviously as they buy a service that is meant to work 100% at any time and from anywhere.

Bad experiences are often caused by network related problems:

• Calls with degraded voice quality
• Interrupted or Dropped calls or VoIP Dropped calls
• Unsuccessful call attempts
• Missed calls that did not ring or calls that were not signaled

Also end device related problems bring negative user experience:

• Empty battery on smartphone
• End device & VoIP Endpoint crashes
• Inability to setup 3 party conference call

For all voice operators, over-the-top (OTT) and legacy networks, these end-device related troubles are very important. Because in the end the customer does not care at all why something is not working! He will blame it on his voice service provider anyway.

As a result when considering the management of customer experience in next generation voice networks, it is very important to look beyond your own network infrastructure. Internet connections, utilization and network problems on the customers’ corporate LAN and end devices cause a large portion of the problems. So it would definitively not be a good idea to consider a service running well, simply because in one’s own network, all systems are up and running.

Instead operators also need to manage:

• End device & VoIP Endpoint firmware versions
• App/softphone software versions
• IP link quality characteristics
• Performance and problems at your Customers’ premises
• Error logs from the end device


Very often problems can be solved by optimizing configurations, codec settings, centralized firmware updates or pro-actively contacting the customer about up-coming problems with his device or software. So a tool is needed that see the endpoint or VoIP device as well as the network. But for all of that, a customer experience management system must provide the full set of information about how the customer truly perceives the service End-to-End.

When looking at customer experience management for VoIP networks, operators should choose a holistic approach and carefully select software solutions that are designed to look far beyond the usual scope and take care about the most important stakeholder in this game – the customer.