We tested our Face Score against 1,566 human-rated faces. Here is exactly what we found.
The RealSmile Face Score reaches a 5-fold cross-validated correlation of r ≈ 0.80 with averaged human attractiveness ratings on the SCUT-FBP5500 benchmark (1,566 faces, each rated by 60 people). Our earlier 17-metric geometry composite did not predict human ratings at all (r ≈ -0.33), so we retired it as a score and now present those metrics only as measurements. To our knowledge, no other consumer face-rating tool publishes any validation of its scoring against human ratings.
Published 2026-07-02 · last reviewed 2026-07-23 · RealSmile
Methodology
The benchmark. SCUT-FBP5500 (Liang et al., 2018) is the standard academic dataset for facial beauty prediction: 5,500 face photos, each independently rated for attractiveness by 60 human raters. The averaged rating for each face is the human ground truth we score against.
What we ran. We put our production pipeline through the benchmark unchanged: the same 68-point facial landmark detector every scan uses, feeding the same scoring model. Nothing was tuned to the benchmark for reporting.
How many faces actually scored. Of the 5,500 benchmark photos, 1,566 produced a complete 68-landmark read from our detector; the rest were rejected by the same detector that rejects a scan on the site. Every figure on this page is measured on those 1,566, not on all 5,500. The surviving set is also not evenly drawn — it is about 79% East Asian and skews young (mean apparent age 25), because detection succeeded more often on some cohorts than others. That is a real limit on how far this number generalises, and it is part of why the in-the-wild figure below is lower.
How we measured it. We used 5-fold cross-validation, so every reported correlation is measured on faces the model never saw during fitting. This is the guardrail against a model that only looks accurate because it memorized its own training data. The headline figure held across the male and female subsets.
1,566
faces scored (of 5,500)
60
human raters per face
5-fold
cross-validation
Results
Face Score (Impression Percentile)
A 128-dimension facial-appearance model of how a 60-person rater panel would place a photo.
r ≈ 0.80
Validated
Face Score, on real user photos
The same model against 46 photos our own users submitted — phone snaps in whatever light they had — rated by real people, median 81 votes each.
r ≈ 0.40
Validated
17-metric geometry composite
The unweighted average of the 17 proportion metrics we previously used as the overall score.
r ≈ -0.33
Retired
The part most tools would hide
Our original overall score was the unweighted average of 17 geometric ratios (symmetry, canthal tilt, facial-width-to-height, golden ratio, and so on). When we checked it against the 1,566 human-rated faces, it showed no positive correlation with how people actually rated those faces (r ≈ -0.33). The individual measurements are real geometry, but blending them into one number did not predict attractiveness. So we stopped calling it a score. Those 17 readings now appear only as a Measurement Map of your proportions, never as a rank. We publish the failure because a claim you can check, including where it broke, is the only kind worth trusting.
Model update — v2
A second-generation kernel model, validated the same way, reaches a 5-fold cross-validated correlation of r ≈ 0.836. v2 is served to Pro members; the r ≈ 0.80 model remains the one behind the one-time report.
How that compares
We are not claiming to be the strongest model in the literature. Peer-reviewed deep-learning models on this same benchmark score higher than we do. The claim that is actually ours: we are the only consumer face-rating tool that publishes a validation you can check at all.
| Tool or model | Publishes validation vs human ratings? | Reported correlation |
|---|---|---|
| RealSmile Face Score | Yes — this page | r ≈ 0.80 (5-fold CV, SCUT-FBP5500) |
| Published deep-learning models (academic benchmark) | Yes — peer-reviewed papers | r ≈ 0.85 to 0.90 |
| Simple geometric / golden-ratio feature models (academic) | Yes — peer-reviewed papers | r ≈ 0.55 to 0.65 |
| QOVES, Umax, PrettyScale (consumer tools) | No published validation found | Not disclosed |
RealSmile Face Score
ValidationYes — this page
rr ≈ 0.80 (5-fold CV, SCUT-FBP5500)
Published deep-learning models (academic benchmark)
ValidationYes — peer-reviewed papers
rr ≈ 0.85 to 0.90
Simple geometric / golden-ratio feature models (academic)
ValidationYes — peer-reviewed papers
rr ≈ 0.55 to 0.65
QOVES, Umax, PrettyScale (consumer tools)
ValidationNo published validation found
rNot disclosed
Scope and limits
The validation covers the Face Score only. It measures agreement with how people rate photos on this benchmark, not dating outcomes, not real-world results. The benchmark also skews toward controlled, front-facing photos.
The Face Score is reported as an impression percentile, where the population average is the 50th percentile by definition. It is never reported as a score out of 100. Your number moves with lighting, angle, and expression, which is exactly why we treat it as feedback on a photo rather than a verdict on a face.
If a competitor publishes their own validation against human ratings, we will link it here.
Cite this
Citation
RealSmile. (2026). Validated Face Score: cross-validated correlation of the RealSmile Impression Percentile against human attractiveness ratings (SCUT-FBP5500). Retrieved from https://realsmile.online/face-score-validation
One-line attribution
“RealSmile's Face Score is validated at r ≈ 0.8 against 1,566 human-rated faces (SCUT-FBP5500, 5-fold cross-validation) — the only consumer face-rating tool to publish a validation of its scoring.”
Canonical URL
https://realsmile.online/face-score-validation
Primary benchmark source: Liang, L., Lin, L., Jin, L., Xie, D., & Li, M. (2018). SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction. International Conference on Pattern Recognition (ICPR). For the full study-to-page index of every source RealSmile cites, see the Research Base.
Measured, not claimed
Everything we know about how accurate this is
Including the numbers that do not flatter us. Every figure below is measured against real scans on this site, carries its sample size, and is dated. Where a claim comes from an external benchmark rather than our own users, it says so — that distinction is the whole point.
Does the score match what people think?
Two different answers, because they are two different questions. A curated benchmark and real phone photos are not the same test, and publishing only the first would be misleading.
r ≈ 0.80 against human ratings
benchmarkn = 1,566 faces · measured 2026-07-01
Five-fold cross-validation of the 128-dimension descriptor model against a published set of human-rated faces. This is the figure most vendors would publish alone.
r ≈ 0.40 against our own raters
our usersn = 46 photos, median 81 votes each · 95% CI 0.12 – 0.62 · measured 2026-08-23
The shipped model scored against real user photos rated by real people in our own pool — phone snaps in whatever light they had, not a curated dataset. Every one of these photos is genuinely unseen: the model was fitted only on the benchmark and has never been trained on a single contributed photo. This is roughly half the benchmark figure, and it is the number that describes what you will actually experience.
About a third of the benchmark-vs-reality gap is range restriction
our usersn = 46 photos · measured 2026-08-23
Correlation shrinks mechanically when a sample is less spread out than the one a model was built on. Benchmark ratings span 1.03–4.67 (SD 0.70); our contributors span 1.62–3.39 (SD 0.42), 1.65x narrower. Corrected to the benchmark’s spread, the in-domain figure is about 0.58 rather than 0.40. We report the uncorrected number as the headline because that is what is measured — but the gap is not all model failure.
The model reaches about 41% of the best score achievable
our usersn = 46 photos · measured 2026-08-23
With a median 81 votes per photo the averaged human label is 0.944 reliable, which caps any model at r ≈ 0.97. Measuring against 1.0 would understate how much room is really left; measured against the attainable ceiling, there is still more than half of it.
A frontier vision model scores r ≈ 0.55 on the same photos
our usersn = 61 photos · measured 2026-08-23
We tested a general-purpose vision model against the same human ratings, two independent passes (self-agreement r = 0.96). On the 23 photos where both scored, ours edged it — 0.48 to 0.40, with overlapping intervals. That 0.48 is a different, smaller subsample than the 46-photo figure above; neither supersedes the other. The honest reading is that neither model is clearly better, and that this task is harder than either number suggests.
How good is the yardstick itself?
A model cannot be more accurate than the human ratings it is measured against. So we measured those too.
Two random people agree on attractiveness at ICC 0.171
our usersn = 78 photos, ~7,000 votes · measured 2026-08-23
One person’s opinion of a face is very nearly worthless as a measurement. This is normal and well documented — it is why a single rating means little.
Averaged over ~100 raters the label is 0.96 reliable
our usersn = median 99 votes/photo · measured 2026-08-23
Averaging destroys the disagreement above. It also sets the ceiling: no model can correlate with a label better than the label correlates with itself, which here is about 0.98.
Above what, exactly?
A percentile is meaningless without naming the crowd it counts against. Ours is a research benchmark of posed studio portraits, which real phone photos beat easily — so every result also shows where you land among people who actually scanned here.
73% of scans rank above the benchmark median
our usersn = 6,507 scans · measured 2026-08-23
The benchmark is posed studio portraits of a mostly young, mostly East Asian cohort. Real phone selfies score systematically high against it — the distribution is nowhere near even (chi-square 1790 on 9 degrees of freedom, where 28 would already be a decisive mismatch). Nothing is wrong with the prediction; the reference group is simply not you. This is why every result shows a second percentile counted against other people who scanned here, where the median scan sits at the 50th rather than the 66th.
The comparison group is re-checked, not assumed
our usersn = 2,025 people · measured 2026-08-23
The "compared with other people who scanned here" table is a snapshot, and a snapshot silently ages as the audience changes. Re-measured against the live population it has drifted by an average of 0.05 percentile points, with a largest single-point move of 1 — still current. One first scan per person, so heavy re-scanners cannot weight the comparison.
How much does the same face move between photos?
This is the number most tools do not publish. It is also the one that matters most to you, because it decides whether a single reading means anything.
The same person on the same day varies by a median of 11 points
our usersn = 1,528 person-days · measured 2026-08-17
A third of people move 20 points or more. The photo moves the score more than the face does.
Two photos on different days cut that spread by about 50%
our usersn = 56 people · measured 2026-08-23
Three photos across different days reach about 53%. Three photos in ONE sitting reach about 18%, and even an unlimited number in one sitting caps out near 28% — roughly half the variance lives between sittings and no amount of repeat-clicking in one session can touch it.
Photo quality does not predict the score
our usersn = 171 people · 95% CI −2.6 – +2.8 points · measured 2026-08-22
We compared each person’s own higher-quality captures against their own lower-quality ones. The difference was 0.11 points — a coin flip. We had been giving lighting and framing advice on the assumption it mattered; our own data does not support it, so we stopped.
How much each measurement can move on its own
Two photos of the same face, same day, minutes apart. A reading has to move more than this before it means anything real. We use these thresholds internally before we will tell you a metric changed — and as you can see, some of them are large.
Between 6 and 31 points, depending on the measurement
The steadiest readings need only a few points of movement before it means anything. The noisiest need far more — those are the ones a change in angle or expression moves as much as a change in your face. We hold every reading to its own threshold before we will tell you it moved, and your report shows the exact figure for each of your measurements.
~3,986 same-day scan pairs per metric · measured 2026-08-23
Change log
Every scan records which version of the measurement engine produced it, so an improvement on our side is never mistaken for a change in your face.
Perfect readings no longer score lower than near-perfect ones. Four measurements — facial symmetry, facial thirds balance, proportion balance and orbital tilt symmetry — are scores where 100 means "perfectly matched" and nothing can go higher. The engine was scoring them by distance from the MIDDLE of their target range, which quietly penalised being at the top of it: a flawless facial-thirds reading scored 42 while a slightly-off one scored 62, and a perfectly symmetric face scored ten points below a slightly asymmetric one. Those four now measure distance from the top of the range instead, so higher genuinely reads as better. If one of these four is among your stronger readings, expect it to go UP; nothing else moves.
Aug 2026
Sharper reference ranges. Several readings — canthal tilt, eye shape, brow-eye proximity, jaw-line slope, lip ratio, jaw taper and brow arch — were being compared against textbook target ranges that turned out to be far wider, or set differently, than real faces measure. That made those readings blunt: they were correct about your face, but too coarse to tell you where you actually sit. All five are now calibrated against 2,071 people who scanned here, so the same measurement now has the precision to place you. One thing to expect: because the scale is finer, your Measurement Map number reads roughly 17 points lower than it did on the previous engine (a typical map score moved from about 66 to about 49). Your face has not changed — the ruler has more marks on it, and where you sit is now a real position rather than a shared top mark.
Aug 2026
Finer measurements. Nose width, midface, eye spacing and proportion balance now keep their decimal precision instead of being rounded to whole numbers, so small real differences between faces show up in the reading rather than being smoothed away.
Aug 2026
Recalibrated canthal tilt and eye shape against a wider reference set. Eye spacing and jaw-line angle moved to measurement-only: both depend on a side profile, and we would rather show you the reading than score it from a front-on photo alone.
Jul 2026
Pose and expression checks. The scan now recognises when a camera angle or a smiling, parted mouth affects a particular reading, and tells you so rather than reporting it as if the shot were ideal.
Jul 2026
Earlier scans, taken before each result recorded which version of the measurement engine produced it. They are kept in full — they are simply compared only against each other.
before Jul 2026
See where your own photo lands on the validated scale.
Get your free Face Score →Free score. Photos are processed in memory, never stored.