Fairness is becoming an increasingly important part of biometric evaluation. For organizations developing, buying or deploying face recognition systems, it is no longer enough to know whether a system is accurate overall. They also need confidence that performance differences across demographic groups have been assessed in a meaningful way.
That sounds straightforward until the assessment begins. There is no single accepted fairness score. Different metrics measure different forms of disparity, and the same system can receive different fairness conclusions depending on the metric selected. For testing laboratories, technology providers and organizations preparing for compliance or certification activities, this creates a practical problem: how can a fairness result be trusted if the choice of metric can change the conclusion?
A recent study conducted by Fime together with Université Caen Normandie, ENSICAEN, CNRS and the GREYC laboratory addresses this question directly. The research compares 19 fairness metrics on four public face datasets and studies three sources of bias: algorithmic choices, variations in data and image quality, and changes in operational decision thresholds. Its main message is clear: fairness cannot be reduced to one universal metric. The metric has to fit the way bias appears in the system.
Why is bias assessment a challenge for certification and conformity evaluation?
Certification and conformity assessment depend on repeatable evidence. A result should be understandable, reproducible and relevant to the conditions in which a product will be used. Fairness assessment adds another layer of complexity because demographic disparity can appear at several points in a biometric system.
It can come from the algorithm itself. It can be influenced by the composition of evaluation data. It can emerge when image quality deteriorates. It can also change when the decision threshold is adjusted to meet a security or user experience objective. A system that looks balanced under one condition may therefore behave differently under another.
This matters for anyone who needs to demonstrate that a biometric system has been evaluated responsibly. If a fairness metric is insensitive to the type of disparity that actually exists, the result may look reassuring without capturing the relevant risk. Conversely, a metric that reacts strongly to one phenomenon may overstate another. The difficulty is not simply calculating a fairness indicator. The difficulty is knowing whether that indicator is the right one for the question being asked.
How can science make fairness assessment more reliable?
The study starts by changing the way fairness metrics are compared. Instead of evaluating each metric in isolation, the researchers place 19 metrics under the same experimental framework and expose them to controlled sources of variation.
For data related and operational bias, the level of disturbance is increased progressively. Image quality is degraded in controlled stages, subgroup representation is reduced in known proportions, and decision thresholds are tightened according to user oriented and security oriented operating regimes. Because the severity order is known, the response of each fairness metric can be checked against a reference.
Two criteria are then used to judge that response. The first is monotonicity: does the metric move consistently as the level of bias increases? The second is stability: does it provide sufficiently consistent conclusions across demographic and intersectional groups? A metric is considered useful for a controlled condition when it satisfies both.
This is an important shift for practical evaluation. Rather than asking whether a metric is mathematically valid in general, the framework asks a more operational question: under which conditions does this metric provide a reliable signal?
Then comes the next problem: with so many fairness metrics, which one should be used?
This is where the problem becomes very concrete. The biometric community now has metrics based on absolute error differences, ratios, inequality measures, score separation, compactness and full score distributions. They are not interchangeable because they do not observe the same thing.
A metric that looks at False Match Rate and False Non Match Rate differences is focused on decision level consequences. A metric that studies score distributions can detect changes before a decision threshold is even applied. Another may be particularly sensitive to the relative difference between the best and worst performing groups.
The study shows that this diversity is not a weakness in itself. The problem appears when a metric is selected without considering the source of bias. In that situation, organizations risk treating the metric as the objective, when it should be a tool chosen to answer a specific assessment question.
What if those metrics disagree?
They may both be telling the truth, but about different aspects of the system.
The experiments show that no single fairness measure provides the most reliable picture in every situation. Some approaches are better at capturing differences in error rates between demographic groups, while others are more sensitive to changes in the underlying biometric score distributions. Other measures become more relevant when the separation between genuine and impostor scores changes, or when the system’s operating threshold is adjusted.
The same applies to image quality. Motion blur, sensor noise, low exposure and JPEG compression do not affect biometric comparisons in the same way. Some degradations mainly alter score separation, while others reshape score distributions or create larger differences in error rates between groups. This means that the most appropriate way to assess fairness can change depending on the type of degradation being considered.
A disagreement between fairness results does not therefore mean that one of them is necessarily wrong. It may simply mean that different aspects of the same system are being measured.
For certification and testing, this distinction is essential. The goal should not be to make every fairness measure produce the same verdict. It should be to demonstrate that the selected approach is appropriate for the type of bias, the risk being assessed and the operating conditions in which the biometric system will be used.
How can this become a practical method for selecting metrics?
The research proposes a structure aware selection principle: start with the expected manifestation of bias, then choose the metric family that is designed to detect it.
If the concern is sampling variability, bounded absolute error metrics can provide a stable view, while distribution metrics can add information about changes in the underlying score distributions. If image degradation changes the separation between genuine and impostor scores, separation-based metrics become relevant. If the main risk comes from the operating threshold, metrics that capture relative inequality or jointly track changes in error rates may be more informative.
The study also supports the use of complementary metrics when one number cannot describe the whole problem. Combining a decision level measure with a score distribution measure can give evaluators two different but useful perspectives on the same system. The important point is that the combination should be motivated by the expected bias mechanism, not applied mechanically.
What should organizations ask before trusting a fairness result?
It changes the conversation. Instead of asking only, “Has this system been tested for fairness?”, organizations should ask how fairness was defined, which demographic groups were considered, under which operating conditions the system was evaluated, which metric was selected and why that metric is relevant to the identified risk.
This approach is also useful when access to the biometric model is limited. The study notes that closed box evaluations can still use group level False Match Rate and False Non Match Rate values with error-based fairness metrics. When full score distributions are available, evaluators can go further and study separation, compactness and distribution changes.
The implications also extend to standards. ISO/IEC 19795-10 defines several aggregated error measures, but the study shows that these variants can behave differently under the same controlled conditions. This suggests that future guidance should not only define how fairness metrics are calculated. It should also help laboratories determine when each metric is appropriate.
From research to meaningful biometric assurance
For Fime, the value of this work lies in turning a complex scientific question into something that can support real evaluation decisions. Fairness is not meaningful because a report contains a score. It becomes meaningful when the measurement is connected to the system, the data, the operating point and the risks of the intended use case.
By contributing to research on the behavior of fairness metrics, Fime is helping build the evidence needed for more robust testing methods and clearer future guidance. This matters for laboratories, technology providers and organizations that need to move from broad commitments on responsible biometrics to assessment methods that can be explained and defended.
The question is no longer simply “Is this face recognition system fair?” The better question is “Have we measured the right form of disparity, with the right metric, under the conditions that matter?”
K. N. Sanon, J. D. Manno, C. Charrier and C. Rosenberger, “When Fairness Metrics Disagree: Operational Implications for Face Recognition Systems,” in IEEE Transactions on Biometrics, Behavior, and Identity Science, doi: 10.1109/TBIOM.2026.3720756.
Download the scientific paper to explore how to select the right fairness metrics for robust and reliable biometric evaluations.