Similarity score
The Similarity score is an aggregate metric that offers a comprehensive assessment of the overall similarity between the real and synthetic data. This score is a reliable indicator of the quality and fidelity of synthetic data generated by the Aindo platform.
Calculation methodology
The Similarity score for a table is derived from the mean of all Bivariate Similarity scores calculated across pairs of variables within the table.
The Bivariate Similarity score measures the distance between two bivariate distributions within a dataset. You can find more details in the section on bivariate distributions. It is computed by subtracting the Total Variation Distance from 1 and then multiplying the result by 100. The Total Variation Distance is defined as one half of the sum of absolute differences between the observed frequencies of each bin in histograms generated from the real and synthetic data.
Interpretation
The Bivariate Similarity score is calculated for every pair of variables in the dataset and it is a number between 0 and 100. A higher Bivariate Similarity score (closer to 100) indicates a greater similarity between the real and synthetic data distributions for the corresponding variables. Conversely, a lower score (closer to 0) suggests weaker similarity, indicating potential disparities in the generated synthetic data compared to the real data.
The Similarity score for a table is the mean of all Bivariate Similarity scores, and therefore also ranges from 0 to 100. When the Similarity score deviates significantly from 100, consult the Univariate and Bivariate distribution comparisons to identify which variables are contributing the most to the discrepancy.