Post

VMAF vs ITU-T P.1203: An empirical comparison for HD videos

VMAF vs ITU-T P.1203: An empirical comparison for HD videos

VMAF [1] has become one of the most useful tools for video quality assessment (VQA) [2]. As a full reference metric, it is widely used on VoD assets, where the sources are easily available. However, VMAF has two drawbacks for live streams: accessing the live source in sync with the encoded output could be impossible, and even when it is possible, it is computationally expensive.

To deal with these issues, several no-reference metrics are available. In this article, we briefly look at one of them, ITU-T Rec. P.1203.1, using its implementation publicly available on GitHub [3], and we compare its scores with VMAF.

ITU-T P.1203 in short

ITU-T Rec. P.1203 is a standard for measuring the Quality of Experience (QoE) of HTTP Adaptive Streaming services. It was trained and validated to predict QoE for transmissions that include initial loading, stalling, and quality variations. Particularly, P.1203.1 mode 3 defines a no-reference metric for HD videos encoded with AVC.

In our tests, we do not take network artifacts (i.e., loading and stalling) into account. For 4K videos or codecs other than AVC, ITU-T Rec. P.1204 should be used instead [4].

Use cases

We typically use VQ metrics to (i) measure the impact of encoding settings on the video streams, (ii) benchmark different encoders, and (iii) define the bitrate ladder for ABR services. Accordingly, we compare P.1203.1 and VMAF only for these three use cases.

Methodology

We are particularly interested in how visual metrics work for live streams. Hence, we use as the source a 60-second sample of a soccer match. Soccer matches contain high-complexity scenes, which makes them a good test of encoder behavior. Additionally, sports events are the most common use case for OTT live services.

We encode the same sample with three encoder implementations: the open source libx264 and two proprietary encoders, which we call Encoder A and Encoder B. Each encoder has vendor-specific features that are not relevant to these tests, since the goal is to analyze the VMAF and P.1203 scores given the same sample and the same encoder. For each encoder, we configure several resolutions and bitrates and measure the visual quality at the output. We introduce neither frame loss nor packet loss, so we only measure artifacts caused by compression.

For VMAF, we use the default model (model/vmaf_v0.6.1.pkl). For P.1203.1, we use mode 3 (i.e., full bitstream analysis). To make both metrics comparable, we normalize VMAF scores to a 0–5 scale (i.e., VMAF/20).

Results

The following results are not intended to determine which model better matches a real Mean Opinion Score (MOS). We only compare the relationship between VMAF and P.1203 scores. Still, these results help to understand the constraints of each model and what to watch for when using them.

VMAF vs ITU-T P.1203

Broadly speaking, P.1203 outputs lower scores than VMAF. The average difference between both scores seems to depend on the encoder implementation. We find the closest match between both metrics with libx264 (see the figure below). This makes sense, given that the P.1203 model was trained with libx264.

VMAF scores plotted against ITU-T P.1203 scores for libx264, Encoder A, and Encoder B

It must be taken into account that VMAF scores could be boosted if the encoder applies picture enhancement techniques [5]. Unfortunately, we cannot confirm whether Encoders A and B apply image enhancement by default, since these are vendor-specific implementations. A new VMAF feature or model could apparently solve this bias in the near future [6].

VMAF and ITU-T P.1203 vs bitrate

A common use of visual quality analysis is defining the bitrate ladder for ABR services, typically by computing the convex hull of the encoder under test. In the next figures, we compare the convex hulls obtained from VMAF and P.1203 scores for each encoder.

VMAF and ITU-T P.1203 scores vs bitrate for libx264

VMAF and ITU-T P.1203 scores vs bitrate for Encoder A

VMAF and ITU-T P.1203 scores vs bitrate for Encoder B

As in the previous section, VMAF scores for Encoder B are much higher than P.1203 scores. Additionally, the convex hulls obtained from both metrics are different.

VMAF and ITU-T P.1203 for encoder benchmarking

Interestingly, the conclusion of an encoder benchmark depends on which metric we use. With VMAF, Encoder B reaches the best bitrate efficiency. With P.1203, Encoder B falls to the bottom of the ranking. The next two figures show this.

VMAF scores vs bitrate for libx264, Encoder A, and Encoder B

ITU-T P.1203 scores vs bitrate for libx264, Encoder A, and Encoder B

Based on our experience with the tested encoders, we believe Encoder B could produce VMAF scores higher than the real MOS due to picture enhancement techniques. However, the real MOS should not be as low as P.1203 predicts.

Discussion

We suggest using either VMAF or P.1203 as a reference metric within the same encoder implementation, and treating their scores as relative rather than absolute values. Said differently, within the same encoder, you can choose which encoding setting gives a better MOS based on either metric, but you still need to validate whether that score meets your quality requirements.

For encoder benchmarking, the state of the art still has some constraints. Nevertheless, we believe VMAF remains the better and more scalable choice for a while: the fix for the inaccuracy caused by picture enhancement seems to be almost ready, and even if it is not enough, you can train your own model independently of the encoder implementations.