We tested our Face Score against 5,500 human-rated faces. Here is exactly what we found.
The RealSmile Face Score reaches a 5-fold cross-validated correlation of r ≈ 0.80 with averaged human attractiveness ratings on the SCUT-FBP5500 benchmark (5,500 faces, each rated by 60 people). Our earlier 17-metric geometry composite did not predict human ratings at all (r ≈ -0.33), so we retired it as a score and now present those metrics only as measurements. To our knowledge, no other consumer face-rating tool publishes any validation of its scoring against human ratings.
Published 2026-07-02 · last reviewed 2026-07-23 · RealSmile
Methodology
The benchmark. SCUT-FBP5500 (Liang et al., 2018) is the standard academic dataset for facial beauty prediction: 5,500 face photos, each independently rated for attractiveness by 60 human raters. The averaged rating for each face is the human ground truth we score against.
What we ran. We put our production pipeline through the benchmark unchanged: the same 68-point facial landmark detector every scan uses, feeding the same scoring model. Nothing was tuned to the benchmark for reporting.
How we measured it. We used 5-fold cross-validation, so every reported correlation is measured on faces the model never saw during fitting. This is the guardrail against a model that only looks accurate because it memorized its own training data. The headline figure held across the male and female subsets.
5,500
benchmark faces
60
human raters per face
5-fold
cross-validation
Results
Face Score (Impression Percentile)
A 128-dimension facial-appearance model of how a 60-person rater panel would place a photo.
r ≈ 0.80
Validated
17-metric geometry composite
The unweighted average of the 17 proportion metrics we previously used as the overall score.
r ≈ -0.33
Retired
The part most tools would hide
Our original overall score was the unweighted average of 17 geometric ratios (symmetry, canthal tilt, facial-width-to-height, golden ratio, and so on). When we checked it against the 5,500 human-rated faces, it showed no positive correlation with how people actually rated those faces (r ≈ -0.33). The individual measurements are real geometry, but blending them into one number did not predict attractiveness. So we stopped calling it a score. Those 17 readings now appear only as a Measurement Map of your proportions, never as a rank. We publish the failure because a claim you can check, including where it broke, is the only kind worth trusting.
Model update — v2
A second-generation kernel model, validated the same way, reaches a 5-fold cross-validated correlation of r ≈ 0.836. v2 is served to Pro members; the r ≈ 0.80 model remains the one behind the one-time report.
How that compares
We are not claiming to be the strongest model in the literature. Peer-reviewed deep-learning models on this same benchmark score higher than we do. The claim that is actually ours: we are the only consumer face-rating tool that publishes a validation you can check at all.
| Tool or model | Publishes validation vs human ratings? | Reported correlation |
|---|---|---|
| RealSmile Face Score | Yes — this page | r ≈ 0.80 (5-fold CV, SCUT-FBP5500) |
| Published deep-learning models (academic benchmark) | Yes — peer-reviewed papers | r ≈ 0.85 to 0.90 |
| Simple geometric / golden-ratio feature models (academic) | Yes — peer-reviewed papers | r ≈ 0.55 to 0.65 |
| QOVES, Umax, PrettyScale (consumer tools) | No published validation found | Not disclosed |
RealSmile Face Score
ValidationYes — this page
rr ≈ 0.80 (5-fold CV, SCUT-FBP5500)
Published deep-learning models (academic benchmark)
ValidationYes — peer-reviewed papers
rr ≈ 0.85 to 0.90
Simple geometric / golden-ratio feature models (academic)
ValidationYes — peer-reviewed papers
rr ≈ 0.55 to 0.65
QOVES, Umax, PrettyScale (consumer tools)
ValidationNo published validation found
rNot disclosed
Scope and limits
The validation covers the Face Score only. It measures agreement with how people rate photos on this benchmark, not dating outcomes, not real-world results. The benchmark also skews toward controlled, front-facing photos.
The Face Score is reported as an impression percentile, where the population average is the 50th percentile by definition. It is never reported as a score out of 100. Your number moves with lighting, angle, and expression, which is exactly why we treat it as feedback on a photo rather than a verdict on a face.
If a competitor publishes their own validation against human ratings, we will link it here.
Cite this
Citation
RealSmile. (2026). Validated Face Score: cross-validated correlation of the RealSmile Impression Percentile against human attractiveness ratings (SCUT-FBP5500). Retrieved from https://realsmile.online/face-score-validation
One-line attribution
“RealSmile's Face Score is validated at r ≈ 0.8 against 5,500 human-rated faces (SCUT-FBP5500, 5-fold cross-validation) — the only consumer face-rating tool to publish a validation of its scoring.”
Canonical URL
https://realsmile.online/face-score-validation
Primary benchmark source: Liang, L., Lin, L., Jin, L., Xie, D., & Li, M. (2018). SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction. International Conference on Pattern Recognition (ICPR). For the full study-to-page index of every source RealSmile cites, see the Research Base.
See where your own photo lands on the validated scale.
Get your free Face Score →Free score. Photos are processed in memory, never stored.